Kimi Pricing 2026 thumbnail featuring a glowing Kimi app logo with the text “Plans, API Costs & Free Tier” on a dark purple and blue neon background.

Kimi Pricing 2026: Plans, API Costs & Free Tier

Kimi is free to start, and that is the fastest way to summarise Moonshot AI’s pricing. The Kimi app gives you unlimited basic chat, file uploads and web search on a $0 Adagio plan, while paid memberships climb from $19 a month to $199 a month for heavy agent and coding use. If you build on the API instead, you pay per token, and the flagship Kimi K3 costs $3 per million input tokens and $15 per million output tokens.

That gap between a free chat app and a metered developer API is the whole story of Kimi pricing, and it trips up most buyers. This guide breaks down every subscription tier, the full per-model API rate card, the real limits of the free plan, and how Kimi’s cost compares against Claude, GPT and DeepSeek. All figures are current as of July 2026 and pulled from Moonshot’s own pricing pages.

The Key Takeaways

  • Free forever tier: the Adagio plan costs $0 with unlimited basic chat, but only about 6 agent credits and no Agent Swarm or Kimi Code.
  • App plans: paid memberships run $19 (Moderato), $39 (Allegretto), $99 (Allegro) and $199 (Vivace) per month.
  • Flagship API: Kimi K3 is $3 / $15 per million input/output tokens, with a cache-hit input rate of just $0.30.
  • Budget API: Kimi K2.6 costs $0.95 / $4 per million tokens, roughly five times cheaper than K3 for output.
  • Cheaper than Western flagships: K3 undercuts Claude Opus 4.8 ($5 / $25) and GPT-5.6 Sol ($5 / $30) on both input and output.

How Much Does Kimi Cost? The Quick Answer

Kimi costs nothing to try. The Kimi app is free on the Adagio tier, and paid plans run from $19 a month up to $199 a month depending on how much agent and coding work you need. On the Moonshot API you pay per token instead of a flat fee, so the flagship Kimi K3 costs $3 per million input and $15 per million output tokens, while the cheaper K2.6 model runs at $0.95 / $4.

Which number matters depends on how you use Kimi. Casual users who chat, upload documents and search the web will mostly live on the free or $19 plan. Developers wiring Kimi into an app, an agent or a coding tool care about the per-token rate, because that is what scales with usage. The rest of this guide covers both sides in full.

Kimi Subscription Plans and Membership Pricing

Moonshot names Kimi’s five membership tiers after musical tempos, from the slow free Adagio plan up to the fast, feature-loaded Vivace tier. Every paid plan is available monthly or annually on Moonshot’s official membership page, and the annual option shaves roughly 20% off the monthly rate. Annual billing on the top plan saves you up to $480 a year.

PlanMonthlyAnnualBest for
AdagioFreeFreeCasual chat and testing
Moderato$19$180/yrSolo users, light agent use
Allegretto$39$372/yrPower users, more coding
Allegro$99$948/yrHeavy agent and Kimi Code
Vivace$199$1,908/yrTeams, maximum concurrency

Adagio (Free)

The free plan gives you unlimited basic chat, file uploads and live web access, which covers most everyday questions. It includes about 6 agent credits, a single concurrent task and 200 professional database calls. Agent Swarm, Kimi Code and Kimi Claw are locked, so serious automation is off the table until you upgrade.

Moderato ($19/mo)

The entry paid tier is the sweet spot for solo users who want real agent work. You get 60 agent credits, two concurrent tasks, 4x agent speed priority and 25 Agent Swarm runs in beta. It also adds Kimi Code at 1x and lifts your database calls to 2,000.

Allegretto ($39/mo)

Allegretto is aimed at power users who lean on Kimi for coding. It bumps you to 150 agent credits, 50 Agent Swarm runs with four concurrent subtasks and Kimi Code at 5x. This is also the first tier to include Kimi Claw, the browser-control agent.

Allegro ($99/mo)

Allegro is the heavy-workload plan, with 360 agent credits, four concurrent tasks and 120 Agent Swarm runs. Kimi Code jumps to 15x and database calls reach 12,000. Most independent developers who run agents daily will not need more than this.

Vivace ($199/mo)

Vivace is the top tier, built for teams and non-stop automation. It delivers 720 agent credits, 240 Agent Swarm runs with eight concurrent subtasks and Kimi Code at 30x. At $199 a month it sits well below what comparable enterprise AI seats cost elsewhere, which is a big part of Kimi’s appeal.

Is Kimi Free? What the Adagio Tier Really Includes

Yes, Kimi is free to use, and that has not changed in 2026. The free Adagio tier still gives you unlimited standard chat, document uploads and web browsing, powered by a capable model rather than a stripped-down one. For asking questions, summarising files and general research, you may never need to pay a cent.

The catch is agentic work. Moonshot has tightened the free tier so that autonomous multi-step tasks, Agent Swarm and Kimi Code sit behind paid plans. You get only a handful of agent credits per period on Adagio, which is fine for a quick test but not for building. If you want Kimi to run long research jobs, control a browser or write and execute code, the $19 Moderato plan is the real starting point. You can read the full model breakdown in our guide to the Kimi K2.6 model.

Kimi API Pricing Per Million Tokens

If you are a developer, the Moonshot API is where Kimi pricing gets interesting. You pay only for the tokens you send and receive, billed per million, and rates vary sharply by model. The newest and most capable model costs the most, while the previous-generation models stay far cheaper. Prompt caching can cut your input cost dramatically when you reuse the same context.

ModelInput /1MOutput /1MCache hitContext
Kimi K3$3.00$15.00$0.301M
Kimi K2.7 Code$0.95$4.00$0.19262K
Kimi K2.6$0.95$4.00n/a262K
Kimi K2.5$0.60$3.00n/a262K
Kimi K2 (legacy)$0.60$2.50$0.15256K

Kimi K3 pricing

The flagship Kimi K3 is a reported 2.8-trillion-parameter model with a 1 million token context window, and it is priced accordingly at $3 input and $15 output per million tokens. The standout number is the cache-hit input rate of just $0.30, a tenth of the miss rate, which rewards workloads that reuse long system prompts or documents. Full specs and benchmarks are in our Kimi K3 deep dive.

Kimi K2.7 Code and K2.6 pricing

Both K2.7 Code and K2.6 cost $0.95 input and $4 output per million tokens, which makes them the value workhorses of the range. K2.7 Code adds a published cache-hit rate of $0.19 and is tuned for long-horizon coding and multi-agent orchestration. For most production apps that do not need a full million tokens of context, this pair is the smart default.

Kimi K2.5 and legacy K2 pricing

The older K2.5 model is cheaper still at $0.60 / $3, and the original Kimi K2 sits at $0.60 / $2.50 with a $0.15 cache rate. These remain available for cost-sensitive workloads or for teams that already built around them. They are slower and less capable than K3, but for high-volume, simple tasks the price difference is hard to ignore.

One extra cost to plan for is tooling. Web-search and certain tool calls are billed on top of token usage, so an agent that searches heavily will run a little above the raw per-token math. Moonshot publishes these fees on its API pricing docs, and they are small per call but add up at scale.

What Does Kimi Actually Cost? A Worked Example

Per-million rates are abstract, so here is a concrete job. Say you send 1 million input tokens and generate 1 million output tokens, which is a large task, roughly a book’s worth of text each way. On Kimi K3 that costs $3 plus $15, so $18 total before any caching. Run the same job on K2.6 and it drops to $0.95 plus $4, or just $4.95.

Caching changes the math again. If most of that input is a repeated document or system prompt that hits the cache, K3’s input cost can fall from $3 toward $0.30 per million, cutting the total closer to $15. For real chat apps, where inputs are short and outputs modest, most requests cost fractions of a cent. The takeaway is simple, pick K3 when you need its reasoning and context, and drop to K2.6 when you do not. Our guide on when to use which AI model walks through those trade-offs.

Kimi Pricing on OpenRouter, Groq and Other Hosts

Moonshot is not the only place to buy Kimi tokens. Third-party inference hosts resell the models, sometimes below Moonshot’s own list price, and they add features like unified billing and faster hardware. On OpenRouter, K3 lists at the same $3 / $15, but effective K2.6 rates can average nearer $0.66 / $3.41 once cheaper providers and caching are blended in.

Groq hosts Kimi K2 for its signature high-speed inference at $1.00 input and $3.00 output, with a $0.50 cached-input rate baked into the price. Groq also stacks a Batch API discount and prompt caching, which together can pull effective costs down to roughly a quarter of the on-demand rate for asynchronous jobs. If throughput and latency matter more than the absolute lowest headline price, a specialist host is often worth it.

Kimi Cost vs Claude, GPT and DeepSeek

Kimi’s pitch has always been frontier-class output at a fraction of Western prices, and the numbers back that up. K3 undercuts both Claude Opus 4.8 and GPT-5.6 Sol on input and output, while offering the same 1 million token context. Only DeepSeek, the other Chinese price-leader, comes in cheaper, though its flagship trades some capability for that discount.

ModelInput /1MOutput /1MNotes
Kimi K3$3.00$15.00Flagship, 1M context
Kimi K2.6$0.95$4.00Cheap workhorse
Claude Opus 4.8$5.00$25.00Frontier coding leader
GPT-5.6 Sol$5.00$30.00OpenAI flagship
DeepSeek V4 Pro$0.44$0.87Cheapest flagship

The honest caveat is capability. On independently verified coding benchmarks, Claude Opus 4.8 still leads the active frontier, and Moonshot’s own agent scores for K3 have yet to be reproduced by third parties. So K3 is not automatically the best model, it is the best-priced model at the top of the context and reasoning range. For a cheaper, capable coder, K2.6 or DeepSeek may serve you better, as covered in our best AI for coding roundup.

Can You Run Kimi for Free by Self-Hosting?

This is where Kimi gets interesting for cost-conscious teams. Moonshot ships its models as open weights, so once released you can download and run them on your own hardware with no per-token fee at all. The catch is that K3’s open weights are expected around late July 2026 rather than live on launch day, and running a 2.8-trillion-parameter model needs serious GPU capacity.

For most people, self-hosting only pays off at very high volume, where the fixed cost of hardware beats a metered API bill. Smaller teams are usually better off with the hosted API or a subscription. The earlier K2 models are already downloadable today, which is why Kimi ranks so highly among the best open-source AI models for anyone who wants full control and zero usage fees.

Which Kimi Plan Should You Choose?

Match the plan to the job. If you just want a smart chatbot for questions, files and web search, the free Adagio tier is enough. If you want Kimi to run agents, browse and code, start at $19 Moderato and climb only when you hit its limits. Developers building products should skip the app entirely and use the API, choosing K3 for hard reasoning and K2.6 for everything else.

If you would rather not manage a Moonshot subscription at all, a multi-model creation app like Fello AI lets you tap Kimi models alongside other frontier models like ChatGPT, Gemini, Grok, Deepseek and more, which is handy when you want to compare outputs. Thanks to Fello AI you can pay only $9.99 a month and use multiple models at once. Whichever route you pick, Kimi remains one of the most aggressively priced frontier options on the market, and the free tier means testing it costs nothing.

FAQ

Is Kimi free to use?

Yes. The free Adagio tier gives you unlimited basic chat, file uploads and web access at no cost. It is capped at roughly 6 agent credits and excludes Agent Swarm, Kimi Code and Kimi Claw, so heavy agent and coding work needs a paid plan starting at $19 a month.

How much does the Kimi API cost per million tokens?

It depends on the model. Flagship Kimi K3 costs $3 per million input tokens and $15 per million output tokens. The cheaper K2.6 and K2.7 Code models cost $0.95 / $4, and older K2.5 runs at $0.60 / $3. Prompt caching can cut input costs by 60 to 80 percent.

What is Kimi K3 pricing?

Kimi K3 costs $3 per million input tokens and $15 per million output tokens, with a cache-hit input rate of just $0.30. It carries a 1 million token context window, so a large 1M-in, 1M-out job costs about $18 before caching.

Is Kimi cheaper than Claude and GPT?

Yes. Kimi K3 at $3 / $15 undercuts both Claude Opus 4.8 ($5 / $25) and GPT-5.6 Sol ($5 / $30) on input and output while matching their 1 million token context. Only DeepSeek is cheaper among flagships, though Claude still leads on verified coding benchmarks.

Can I run Kimi for free by self-hosting?

Eventually, yes. Moonshot releases its models as open weights, so you can run them locally with no per-token fee. Earlier K2 models are downloadable now, while K3’s open weights are expected around late July 2026. Self-hosting only pays off at high volume, since a 2.8-trillion-parameter model needs heavy GPU hardware.

Share Now!

Facebook
X
LinkedIn
Threads
Email

Get Exclusive AI Tips to Your Inbox!

Stay ahead with expert AI insights trusted by top tech professionals!