Skip to content

Kimi K3 and K2.7 Code API Pricing Explained

Current Kimi K3 and K2.7 Code API pricing, cache discounts, context limits, and how to choose between the flagship and coding-specific model.

MGMCSA Guru Team August 15, 2026 3 min read
Kimi K3 and K2.7 Code API pricing comparison

Moonshot’s current API lineup has two useful choices for developers. Kimi K3 is the flagship model with native vision and a 1-million-token context window. Kimi K2.7 Code is a coding-specific model with a 256K context window and lower rates.

Current pricing snapshot

Moonshot’s international K3 announcement lists these prices per million tokens:

Kimi K3 international API pricing

Cached input $0.30 per million tokens
Uncached input $3.00 per million tokens
Output $15.00 per million tokens

Moonshot’s regional platform lists K2.7 Code at a substantially lower rate than K3. Because the regional page bills in yuan and international gateways may use different prices, check the account and endpoint you will actually call before converting the figures into a budget.

K3 or K2.7 Code

Current Kimi choices for developers

kimi-k3 Flagship model, native vision, 1M context, low/high/max reasoning effort
kimi-k2.7-code Coding-specific model, 256K context, lower API cost

K3 fits tasks that need images, a very large context, or the strongest current Kimi model. K2.7 Code is easier to justify for routine repository work where text and 256K context are enough.

The two model IDs are not interchangeable. Kimi Code subscriptions may also use service aliases such as kimi-for-coding, which are not the same as pay-as-you-go API model IDs.

Cache discounts

Kimi charges much less for cached input than for fresh input. Coding agents repeatedly send system instructions, repository context, and conversation history, so a high cache-hit rate can change the total cost substantially.

Do not estimate a whole session using only the cached rate. The first request, changed files, and new conversation content still count as uncached input. Output is also the most expensive part of a K3 request, so long reasoning traces and verbose answers can dominate the bill.

Estimate a coding session

Record cached input, uncached input, and output separately. Multiply each token count by its per-million rate, then add the three results. Repeat the calculation with a representative week of actual usage rather than one short prompt.

For example, a workflow with frequent small edits may benefit heavily from caching. A one-off repository analysis may send far more uncached input and produce a longer answer. Those sessions can use the same model but have very different costs.

Kimi pricing check

  • Choose K3 or K2.7 Code based on the task
  • Check pricing for the same region and endpoint as the API key
  • Separate cached input, uncached input, and output
  • Measure a representative session before setting a monthly budget
  • Recheck the live model page when Moonshot releases a new version

K3 is no longer a cheap drop-in replacement for every task. It is the flagship option for long-context and multimodal work. K2.7 Code remains the practical lower-cost model for text-based coding sessions.

For setup, see Kimi K3 with Claude Code, Kimi K3 with OpenCode, and K2.7 Code with Aider.

Frequently asked questions

How much does Kimi K3 cost?

Moonshot's international announcement lists K3 at $0.30 per million cached input tokens, $3 per million uncached input tokens, and $15 per million output tokens.

Which model is cheaper for coding?

K2.7 Code is the lower-cost coding-specific option. K3 costs more but adds the flagship reasoning, native vision, and 1M-token context.

Why are Kimi prices different on another site?

Regional Moonshot platforms use different currencies, and gateways may set their own rates or discounts. Check the price for the endpoint and account you actually use.

Sources & further reading

Official vendor documentation referenced while writing this guide.

MG

MCSA Guru Team

IT & Systems Administration

We are working IT pros and system administrators who spend our days in Windows Server, Microsoft 365, and the wider Microsoft stack. MCSA Guru is where we write down the fixes and walkthroughs we wish we had found the first time.

MCSA Guru provides independent, educational IT guidance. Microsoft, Windows, Windows Server, Microsoft 365, Exchange, and Microsoft Teams are trademarks of Microsoft Corporation; Docker is a trademark of Docker, Inc. MCSA Guru is not affiliated with or endorsed by Microsoft or Docker. Always test changes in a safe environment before applying them in production.

Related guides

Fixing something right now?

Jump straight into the guide library or search for the exact error or task you are dealing with.