Kimi (Moonshot AI)

Kimi is Moonshot AI's model family. TokenSense routes supported native kimi- models through Chat Completions. This page covers setup, supported models, and how cache-hit billing works.

Connect Kimi to TokenSense

Step 1: Go to your dashboard → Settings → Providers → add your Moonshot AI API key.

Step 2: Point your workflow at your TokenSense endpoint and use any kimi- model name. TokenSense recognises the prefix and routes to Moonshot automatically — there is nothing else to configure.

Step 3: Authenticate with your TokenSense API key. Costs appear in your dashboard within seconds, attributed to the workflow that made the call.

OpenAI-compatible. Kimi speaks the OpenAI chat format, so any tool that can call OpenAI can call Kimi through TokenSense by changing the model name. No request translation, no separate SDK.

Supported models

Current native catalog choices include kimi-k3, kimi-k2.7-code, kimi-k2.7-code-highspeed, and kimi-k2.6. Highspeed uses a verified manual catalog entry because the native feed omits it. Check current rates in your dashboard and Moonshot's pricing before comparing costs.

How cache-hit billing works

Kimi's cache discount varies by model rather than being a single provider-wide multiplier, so TokenSense stores the flat cache-hit input rate for each model. When Moonshot reports cached tokens on a response, those tokens are billed at that rate and the rest at the standard input rate. Your dashboard shows the cached token count alongside the regular ones, so you can see what the cache actually saved you.

One limitation worth knowing: Moonshot's documented streaming schema does not include a cached-token count. When you stream a response and no cached count arrives on the final chunk, TokenSense bills the full input rate rather than guessing. Non-streaming estimates use the usage buckets returned by Moonshot; reconcile them against your provider bill when checking cache savings. If cache savings matter to your workload, prefer non-streaming requests for the calls where it counts.

What Kimi does not support

Kimi is a chat and reasoning family only. Through TokenSense it works with /v1/chat/completions. It has no embeddings, image generation, or audio models, so requests to /v1/embeddings, /v1/images/generations, and the audio endpoints need a different provider. See the provider overview for which providers cover which modality.

Legacy model names. The older moonshot-v1-* model names are not part of the TokenSense catalog. Use the kimi- names above.

Using Kimi in your workflows

n8n: Select a kimi- model in the TokenSense node, or set your OpenAI-compatible node's model field to a Kimi model name.

Make / Zapier: Use an HTTP module pointed at your TokenSense endpoint with a kimi- model in the body. See the Make & Zapier guide.

Custom code: Point your OpenAI client's base URL at your TokenSense endpoint and set the model to a kimi- name. See the API reference.

Related