This page covers the models Kimi Code provides and how to switch between them in each client.
Kimi Code currently offers two models—Kimi K3 and Kimi K2.7 Code—across four model IDs, selectable by model ID in clients or third-party tools. Model specs:
Recommended model launch
k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k. Ideal for everyday Q&A, code completion, routine feature development, and single-file or small-file edits — video input is not supported.
💓 Reminder
Switching from K3 (1M) to K3-256k: When switching from
k3(1M) tok3-256k, if the current session's context already exceeds 256k, some coding tools such as Kimi Code CLI and Claude Code will perform a compact on the tool side.Switching recommendations:
Switching from K3-256k to K3 (1M): When switching from
k3-256ktok3(1M), ifk3-256kis close to the 256k limit and you don't want compact to lose information, you can switch directly to 1M. The current version switching from 256k to 1M does not affect the cache.
| Model ID | k3 | k3-256k | kimi-for-coding | kimi-for-coding-highspeed |
|---|---|---|---|---|
| Model version | Kimi K3 | Kimi K3 | Kimi K2.7 Code | K2.7 Code HighSpeed |
| Description | Kimi's most capable flagship coding model: 2.8T parameters, 1M context window | The 256K context version of Kimi K3, effectively reducing consumption | Good at code completion and routine development tasks | The high-speed version of K2.7 Code, with the same coding ability and ~5–6× faster output |
| Speed | Regular | Regular | Regular | HighSpeed (6× speed, 3× quota usage) |
| Context window | Up to 1M (for higher-tier members) | 256k only | 256k | 256k |
| Reasoning | reasoning_effort:low / high / max | reasoning_effort:low / high / max (default high) | Thinking:ON | Thinking:ON |
| Availability | Available to Moderato and above; 1M context for Allegretto and above | Available to all Moderato members and above | All members | Allegretto plan or above |
| Multimodal input | Image, video | Image only | Image, video | Image, video |
Need a higher membership plan?
Different membership plans unlock different models, context windows, and speeds. Upgrade your plan →
After switching models, the context cache built earlier no longer hits on the new model, so that context has to be re-prefilled. Usage therefore looks higher right after switching. Recommended action:
When the requested capability exceeds your plan's entitlements, the server returns 401. Three common cases:
k3, k3-256k — upgrade to Moderato or above.k3 supports up to 256K context; up to 1M context is available on Allegretto and higher tiers. k3-256k has a fixed 256K context limit.kimi-for-coding-highspeed.For the full error text and how to handle it, see the Error Reference.
Two common reasons:
kimi-for-coding-highspeed; a wrong value silently falls back to the standard kimi-for-coding — no error, no speedup.Switching reasoning effort invalidates the context cache you've built up, so context that would have hit the cache must be re-prefilled. To avoid triggering re-prefill too often:
Usage notes
k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeed). Entering a model version name like Kimi K3 or K2.7 Code will cause the call to fail.Ways to switch to the target model:
/model to switch models—no config changes needed; if the latest model isn't listed yet, /logout and sign in again with /login.Set the tool's Model ID to the target model. Detailed steps:
Kimi Code API supports both OpenAI and Anthropic protocols. Base URLs:
| Protocol | Base URL |
|---|---|
| OpenAI compatible | https://api.kimi.com/coding/v1 |
| Anthropic compatible | https://api.kimi.com/coding/ |
For detailed setup steps, see the corresponding tool guide:
Before using K3 in third-party tools
K3's setup differs slightly from K2.7 Code. Before using it, check the two points below:
1048576 to use K3's full up-to-1M context.low / high / max; the effort a tool sends is mapped as below:# default
null / undefined → high
any other unknown → HTTP 400 error
# → max
ultra / max / xhigh → max
# → high (recommended)
high / medium → high
# → low
low / minimum / light → low
# → thinking disabled
none → thinking.type disabled