Moonshot AI published the full Kimi K3 weights on July 27, 2026, eleven days after the model went live in the Kimi app. The Hugging Face repository carries 96 safetensors shards totalling roughly 1.56 TB. At 2.8 trillion total parameters, with 104 billion active per token, K3 is the largest open model anyone has released. Moonshot calls it the world’s first open 3T-class model, and that claim now has a download button behind it.

The independent picture has moved a long way since launch, and not in K3’s favour. Kimi K3 now scores 43.6 on the Artificial Analysis Intelligence Index v4.3.2, which makes it third among open-weight models rather than first, behind Xiaomi’s MiMo-V2.6-Pro at 46.3 and Z.ai’s GLM-5.3 at 44.8. The boards humans vote on say the opposite, and K3 still leads the open field on both of them. Below you get the verified specs, what the licence actually permits, the official API rates, and an honest read on how K3 stacks up against Claude Opus 5.5, which replaced Opus 5 as Anthropic’s default on September 22, 2026.

The Key Takeaways

  • The open weights shipped on July 27, 2026, 96 shards and about 1.56 TB on Hugging Face, under a custom Kimi K3 License rather than MIT.
  • A Mixture-of-Experts design with 2.8 trillion total parameters, 104 billion activated per token, and a 1,048,576-token context window.
  • Accepts text, image, and video input, always reasons, and now exposes low, high and max reasoning effort rather than max only.
  • Official API pricing is $3.00 per million input tokens and $15.00 output, dropping to $0.30 on cached input, with the cache write itself billed at $3.00 or $6.00 depending on how long the entry lives.
  • Scores 43.6 on Artificial Analysis Intelligence Index v4.3.2, third among open weights behind MiMo-V2.6-Pro at 46.3 and GLM-5.3 at 44.8, but still the best-placed open model on the boards humans vote on.

What is Kimi K3?

From the publisher

Every AI model in one app

Fello AI puts GPT-6, Claude 5.5, Gemini 3.8, Grok 4.7 and more in one native Mac and iPhone app.

Download now!

Kimi K3 is Moonshot AI’s flagship large language model, launched on July 16, 2026 and released as open weights on July 27. It uses a Mixture-of-Experts design with 2.8 trillion total parameters, of which 104 billion activate on any given token, and it reads a 1-million-token context window. Moonshot’s own documentation confirms it accepts text, image and video input.

The gap between launch and weights was planned, and Moonshot said so at the time. Its launch post promised the files “by July 27, 2026” while the team worked “closely with inference partners and open-source maintainers to align technical details and ensure a reliable rollout across the ecosystem”. The technical report shipped with the weights rather than trailing them.

SpecKimi K3
DeveloperMoonshot AI
ReleasedJuly 16, 2026 (weights July 27, 2026)
ArchitectureMixture-of-Experts (MoE)
Total parameters2.8 trillion
Activated parameters104 billion per token
Layers93 (69 Kimi Delta Attention, 24 Gated MLA)
Experts896 total, 16 selected per token
Vision encoderMoonViT-V2, 401M parameters
Context window1,048,576 tokens
Input typesText, image, video
QuantizationMXFP4 weights, MXFP8 activations
Reasoning controlreasoning_effort low / high / max, default max
LicenceKimi K3 License (custom, not MIT)

At 2.8 trillion parameters, Kimi K3 is still the largest open model released to date, although the margin has narrowed since July: Alibaba’s Qwen3.8 2.4T A95B publishes 2.45 trillion, where the nearest rival used to be DeepSeek’s V4-Pro at 1.6 trillion. The Mixture-of-Experts design means only a small fraction of those parameters fire on any given token, so the effective compute cost stays far below what the raw count suggests. Activating 16 of 896 experts is an unusually high sparsity ratio. That choice is central to how Moonshot squeezed frontier performance out of the architecture.

Kimi K3 Intelligence, Performance & Price Analysis by Artificial Analysis [source]

Kimi K3 Open Weights: What Actually Shipped

This is the part that changed on July 27, and it is the reason K3 matters more than its benchmark line. Every other model at this capability level is an API you rent. K3 is a file you can hold.

What is in the repository

The Hugging Face repo contains 120 files, of which 96 are safetensors shards, and the whole download comes to roughly 1.56 TB. Moonshot ships the model in native MXFP4 weights with MXFP8 activations, applied through quantization-aware training from the supervised fine-tuning stage onward. That is why a 2.8-trillion-parameter model fits in 1.56 TB instead of several times that, and it means you are not downloading a lossy community quant, you are downloading what Moonshot trained.

The pull was immediate and it did not stop. A day after the files went up, the repository showed roughly 99,000 downloads and more than 7,300 likes, counted on July 28, 2026. By September 24 Hugging Face was reporting 1.86 million downloads over the trailing month and 11,500 likes. For a checkpoint almost nobody can host on their own machine, that is a lot of traffic, and much of it will be inference providers, quantization projects and research labs rather than individuals.

Moonshot names three inference engines as recommended runtimes, vLLM, SGLang and TokenSpeed, each with its own published recipe for K3. The technical report shipped with the weights rather than trailing them, so the architecture details are public as well.

The Kimi K3 License is not MIT

Read this before you plan a product around it.

The licence grants the usual rights to use, copy, modify, distribute, sublicense, fine-tune and sell, which reads like MIT for most readers. Two clauses then attach conditions that MIT does not have. Tencent took the opposite route on August 28, 2026, publishing Hy4 Preview’s 770B weights under plain Apache 2.0.

The first is a revenue gate on reselling. If you run a model-as-a-service business and your revenue passes $20 million over any consecutive 12 months, you must sign a separate agreement with Moonshot before using K3 commercially. The second is an attribution rule. Any product built on K3 with more than 100 million monthly active users, or more than $20 million in monthly revenue, has to display the name “Kimi K3” prominently in its interface.

Both conditions are waived for purely internal use, meaning anything that never exposes the model or its outputs to third parties. For an individual developer, a startup or a research group, the practical effect is nil. For a hyperscaler reselling inference, it is a negotiation. That trade used to be the whole argument for K3, top capability in exchange for a licence you had to read properly. It is a weaker argument now, because two open models score above K3 on the Artificial Analysis index and one of them, Xiaomi’s MiMo-V2.6-Pro, ships under plain MIT. The other is Z.ai’s GLM line, where the flagship GLM-5.3 carries its own bespoke licence and only GLM 5.3 Flash is MIT.

What it takes to run Kimi K3 yourself

Downloadable is not the same as runnable. A 1.56 TB checkpoint needs a multi-node GPU cluster to serve at full context, so “you can self-host it” is true for labs and serious infrastructure teams, not for someone with a single workstation. If you want an open model you can actually run on your own hardware, the smaller Kimi models and GLM remain the realistic options. Our roundup of the best open source AI models ranks them by what the hardware demands.

The Architecture Behind the Jump

Moonshot reports roughly a 2.5x improvement in overall scaling efficiency compared to Kimi K2. In plain terms, K3 converts raw compute into usable intelligence far more effectively than the previous generation. Three changes carry most of the weight, and all three are now documented in the public technical report.

Kimi Delta Attention

Kimi Delta Attention (KDA) is a hybrid linear attention mechanism, and 69 of the model’s 93 layers use it, with the remaining 24 running Gated MLA. Standard attention gets slower and more memory-hungry as sequences grow, which is exactly the problem at a million tokens. Moonshot reports KDA enables up to 6.3x faster decoding in million-token contexts, which matters for agentic work where a model repeatedly reads and reasons over huge stretches of code or documents.

Attention Residuals

Attention Residuals (AttnRes) address how information moves through the depth of the network rather than across sequence length. Instead of accumulating representations uniformly layer by layer, AttnRes selectively retrieves representations across depth. Moonshot says this delivers about 25% higher training efficiency while adding less than 2% in extra compute, and it open-sourced the technique earlier in 2026 before folding it into K3.

Stable LatentMoE and MXFP4 training

The third piece is the Stable LatentMoE framework, which makes the aggressive 16-of-896 expert sparsity trainable without the instability that usually comes with pushing MoE that far. Sitting alongside it is the decision to train in MXFP4 from the fine-tuning stage rather than quantising afterwards, which is what keeps the released weights both small and faithful to the model Moonshot evaluated.

One demonstration Moonshot highlighted is telling. The team pointed K3 at optimizing its own AttnRes training kernel at production scale, 96 layers, an 8,192-dimension model and 8,192 tokens. Over 15 hours of nonstop iteration, K3 designed a novel two-phase kernel algorithm, fused kernels while preserving numerics, and cut forward-plus-backward time from 283.6 ms to 114.4 ms. That is the kind of long-horizon engineering task most models stall on within minutes.

A 1-million-token context window

Kimi K3 handles up to 1,048,576 tokens in a single context, roughly four times the 256K window of K2.6. That is enough to hold an entire codebase, a stack of long documents, or a lengthy multi-step agent run without losing the thread. It also matches the 1M window Anthropic documents for Claude Opus 5.5, Fable 5.1 and Sonnet 5.5 alike, so the long-context advantage Kimi used to hold over the closed flagships has narrowed to a tie.

Multimodal input and always-on reasoning

K3 accepts text, images, and video, which pushes it past the text-first K2 line into proper multimodal territory. Moonshot’s own K3 quickstart documents video input through the files API, and the model card reports 90.0 on Video-MME with subtitles. Reasoning is always on, and the reasoning_effort parameter now supports low, high and max, with max as the default. At launch only max was available, so the promised effort levels have since arrived.

How Kimi K3 Performs on Benchmarks

Two separate questions live here. What do the independent boards say about Kimi K3, and what does Moonshot’s own evaluation suite claim? The independent picture only arrived after launch, so it is the better starting point, and the vendor numbers make more sense once you have seen it.

What the independent boards say

Artificial Analysis now scores Kimi K3 at 43.6 on Intelligence Index v4.3.2, read off its model page on September 24, 2026. Most of the drop from the 57 this article first reported is a rescale rather than a regression, because the index moved through v4.2 and v4.3 in September and every score on the board came down with it. What genuinely changed is the company K3 keeps. Claude Opus 5.5 sits at 57.6, Claude Fable 5.1 at 53.4 and GPT-6 Astra at 52.7, and among open weights K3 is now third, behind MiMo-V2.6-Pro at 46.3 and GLM-5.3 at 44.8.

Two of the numbers this article used to quote no longer exist. Artificial Analysis folded its Agentic Index into the main index and then retired its Coding Index outright, so neither composite can be refreshed. What survives is speed and cost, and both still cut against K3. Artificial Analysis clocks it at about 37 tokens per second, against 54 for Claude Opus 5 and 116 for GPT-6 Sol, which makes it the slowest model in the frontier group. It charges $2.00 per Intelligence Index task, against $5.98 for Opus 5.5 and $0.13 for MiMo-V2.6-Pro.

On Arena’s WebDev leaderboard, where humans blind-vote on generated front ends, K3 now sits at 1,659.6 from 13,252 votes, read on September 24, 2026. That is ninth overall and first among open-weight models. All eight entries above it are closed: Claude Opus 5.5 at 1,818.4, GPT-6 Astra at 1,792.2, Claude Fable 5.1 at 1,754.7, two Claude Opus 5 effort settings, GPT-6 Sol at 1,685.9 and two Qwen3.8 Max snapshots. K3 held the top spot outright at launch, so this is a board it has been overtaken on, not one it fell off.

Arena’s Agent board, which rates how well models orchestrate tools on real tasks, now publishes an overall rank, and it places Kimi K3 eighth at +0.062, top of the open field and ahead of Tencent’s Hy4 Preview, DeepSeek V4.1 Flash and both GLM entries. The five underlying signals are worth reading separately: K3 is fifth on task completion, the one that says whether the job actually got done, but seventeenth on steerability and nineteenth on recovering from a failed shell command. On the general text board it sits at number 17 of 125 with 1,484.8 points from 20,987 votes. MiMo-V2.6-Pro, the model that now outscores it on the index, has no Arena votes at all.

Moonshot’s own coding benchmarks

Moonshot published this suite at launch in July 2026, comparing Kimi K3 against Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5 and GLM-5.2, all run at maximum thinking effort. Read it as a July snapshot: every closed model in it has since been superseded, and the Terminal-Bench row is version 2.1, which the industry has replaced with the far harder 4.0. These are also the vendor’s own runs, with K3 evaluated on the Kimi Code harness in several rows, so treat them as Moonshot’s claims rather than independent results.

BenchmarkKimi K3Fable 5GPT-5.6 SolOpus 4.8GPT-5.5GLM-5.2
DeepSWE67.570.073.059.067.046.2
ProgramBench77.876.877.671.970.863.7
Terminal-Bench 2.188.388.088.884.683.482.7
FrontierSWE81.286.671.366.764.967.3
SWE-Marathon42.035.039.040.014.013.0
Kimi Code Bench 2.072.976.964.871.769.064.2

Reading this straight, Kimi K3 wins ProgramBench and SWE-Marathon outright and sits within half a point of the leader on Terminal-Bench 2.1. On FrontierSWE it lands second behind Fable 5 but well ahead of everything else. Moonshot’s own footnotes add a caveat worth carrying, Claude Fable 5 hit safety fallbacks on 35% of the SWE-Marathon tasks in that evaluation, which may have dragged its score down.

Agentic and knowledge work benchmarks

BenchmarkKimi K3Fable 5GPT-5.6 SolOpus 4.8GLM-5.2
BrowseComp91.288.090.484.3n/a
MCPMark-Verified94.587.492.976.4n/a
GDPval-AA v2 (Elo)1,6861,7471,7361,5931,510
AA-Briefcase (Elo)1,5481,5831,4951,3541,260
JobBench54.357.445.448.443.4
SpreadsheetBench 234.834.732.431.628.1
AutomationBench30.829.129.727.212.9
OSWorld 2.058.366.162.655.7n/a

Kimi K3 leads BrowseComp, MCPMark, SpreadsheetBench 2 and AutomationBench, and takes second on AA-Briefcase, the long-horizon agentic benchmark, beating GPT-5.6 Sol and trailing only Fable 5. One footnote matters on BrowseComp, that 91.2 uses a context-compaction strategy triggered at 300K tokens. Run with the full 1M window and no context management, K3 scores 90.4, which is still the best number in the column.

On computer use, Fable 5 and GPT-5.6 Sol both beat K3 on OSWorld 2.0, so the open model is not ahead everywhere. It posted 93.5% on GPQA Diamond, level with GPT-5.5 and ahead of GLM-5.2’s 91.2, and 94.6 on Harvey Lab-AA, the best score in that row from any model in the table.

Moonshot’s own summary is more restrained than the coverage was. Its launch blog says K3’s overall performance “still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol”, while demonstrating frontier-level performance across the suite. That sentence is still on the page two months later, and both models it names have been replaced since, by Claude Fable 5.1 and GPT-6 Sol. When a lab undersells its own model, that is usually the number to trust.

Kimi K3 vs Claude Opus 5.5

The comparison everyone runs is Kimi K3 vs Claude, and the Claude side has moved twice since K3 launched. Anthropic shipped Claude Opus 5 on July 24, which is why Moonshot’s launch table above stays attributed to Opus 4.8. Then on September 22 it shipped Claude Opus 5.5 and moved Opus 5 to its Legacy list. For a decision you are making today, Opus 5.5 is the model on the other side of the table, and it is both cheaper and stronger than the Opus 5 this comparison was originally written against.

ModelDeveloperContextOpen weightsAPI price per 1M
Kimi K3Moonshot AI1,048,576Yes, Kimi K3 License$3 / $15
Claude Opus 5.5Anthropic1MNo$4 / $20
Claude Opus 5 (Legacy)Anthropic1MNo$5 / $25
Claude Fable 5.1Anthropic1MNo$10 / $50
MiMo-V2.6-ProXiaomi1MYes, MIT$0.435 / $0.87
GLM-5.3Z.ai1MYes, GLM 5.3 License$1.40 / $4.40
Kimi K2.6Moonshot AI262,144Yes, modified MIT$0.95 / $4

On price K3 still wins, by less than it used to: $3 / $15 against $4 / $20 for Opus 5.5, which is 25% cheaper where it undercut Opus 5 by 40%, and against $10 / $50 for Fable 5.1. On measured capability the gap went the other way, from four points to fourteen on the Artificial Analysis Intelligence Index, and Opus 5.5 leads Arena’s WebDev board by 159 points. Anthropic publishes no SWE-bench figure for either Opus, so anyone quoting one to you is not quoting Anthropic.

The real split is not the score, it is what you get to own. Opus 5.5 is the stronger model and the safer default for production coding, and Anthropic’s own documentation now tells developers to start with it for most workloads. K3 is the one you can download, audit, fine-tune and run inside your own network, which no Claude model allows at any price. If you are choosing between flagships for a specific job, our guide on when to use which AI model breaks the decision down task by task.

Kimi K3 pricing

Moonshot has now published official rates, which retires the third-party estimates that circulated at launch. Kimi K3 pricing is $3.00 per million input tokens and $15.00 per million output tokens, with cached input at $0.30, a 90% discount that matters a lot if you keep re-sending the same long context. K3 also bills the cache write itself, $3.00 for a five-minute entry or $6.00 for a one-hour one, which the K2 models do not.

ModelInput (cache miss)Input (cache hit)OutputContext
kimi-k3$3.00$0.30$15.001,048,576
kimi-k2.7-code$0.95$0.19$4.00262,144
kimi-k2.7-code-highspeed$1.90$0.38$8.00262,144
kimi-k2.6$0.95$0.16$4.00262,144

Those rates no longer match Claude Sonnet 5.5, which runs $2 in and $10 out, so K3 is 50% dearer than Sonnet on both halves of the bill, and since Sonnet 5.5 launched on 28 September it also scores higher than K3, 56.0 against 43.6 on the Artificial Analysis Intelligence Index. Against the other two Claude models above it, K3 still runs 25% cheaper than Opus 5.5 and 70% cheaper than Fable 5.1. Web search has changed shape as well. Moonshot moved it to standalone REST endpoints in September and bills only successful calls, at $0.002 for Web Search Basic, $0.003 for Web Search Pro and $0.002 for URL Fetch.

K3 is a clear step up in cost from the K2 line, which reflects its bigger footprint and always-on reasoning. If your workload does not need frontier-level scale, K2.6 or K2.7 Code remain much cheaper, and kimi-k2.5 is no longer an option: Moonshot discontinued it and the whole moonshot-v1 series on August 31, 2026, and calls to them now return a 404. On the API platform, K3 becomes available once you have made a successful top-up of at least $1, a cumulative $5 earns a $5 voucher, and your cumulative top-up sets your rate limits. Our full Kimi pricing breakdown compares every plan and API rate side by side.

Demand outran Moonshot’s hardware almost immediately. On July 19, 2026, three days after launch, the company paused new Kimi subscriptions, saying demand over the previous 48 hours had “pushed close to the limits of our current capacity” and that it would “reopen new subscription spots in batches” as it added hardware, in a statement carried by the South China Morning Post. Existing subscribers kept their access, and the same post promised to split membership into Kimi Membership for the web, app and Work, and Kimi Code Membership for coding. The launch top-up rebate that Moonshot cut short for the same reason has gone from the pricing docs. What the docs record instead, dated September 2026, is a billing change: usage is now deducted at a flat 50% cash and 50% vouchers.

How to use Kimi K3

There are four ways to reach Kimi K3 now that the weights are public, and which one fits depends on whether you want a chat window, a terminal agent, an API or the raw files. Here is the practical order.

  1. Kimi app and Playground. The fastest route is the Kimi app or the web Playground, where K3 is selectable for logged-in users. This is the no-code option for testing the model on your own prompts.
  2. Kimi Code CLI. Moonshot says K3 works best inside its own terminal agent, where you pick the model with the /model command. K3 also plugs into Codex CLI, Claude Code, OpenCode and Hermes Agent through Moonshot’s documented integrations.
  3. The API. Developers call kimi-k3 through the Kimi API platform, which offers OpenAI-compatible and Anthropic-compatible endpoints. It supports tool calling, JSON-schema structured output, context caching and the reasoning_effort control, and K3 is also listed on aggregators like OpenRouter.
  4. Open weights. The full checkpoint is on Hugging Face under the Kimi K3 License, and Moonshot recommends vLLM, SGLang or TokenSpeed to serve it. Budget for the 1.56 TB download and the multi-node hardware it needs.

Would you rather not juggle a separate subscription for every new model? Apps like Fello put powerful current AI models in one place, which helps while a launch like this one settles. For a wider view of the open field, our roundup of the best open-source AI models available right now puts K3’s lineage in context. If you are weighing Chinese open models specifically, our explainer on what GLM is and how it competes is a useful companion.

The Business Behind Moonshot AI

Kimi K3 arrives at a pivotal moment for Moonshot AI, the Beijing startup founded in 2023 by Yang Zhilin. In July the company was reported to be raising between $1 billion and $2 billion at a valuation of up to $31.5 billion, roughly 50% above the $20 billion it reached in May and around seven times what it was worth in December 2025. That number has already been overtaken. Reuters reported on September 3, 2026 that Moonshot had confidentially filed for a Hong Kong listing, aiming to raise $3 billion at a valuation of $50 billion in an ongoing round, working with Goldman Sachs, CICC and Deutsche Bank. That is the listing TechNode had flagged in July as possible within six months, and Moonshot declined to comment on the filing.

The open-source strategy is central to the momentum. When Moonshot open-sourced Kimi K2.6 in April 2026, it climbed to become the second most-used large language model on OpenRouter by mid-2026. That put a lab founded three years earlier ahead of most Western frontier labs on independent usage leaderboards. Releasing K3’s weights extends that playbook to a frontier-class model.

The timing was not accidental. Kimi K3 landed just before the 2026 World Artificial Intelligence Conference in Shanghai, where Chinese President Xi Jinping was expected to outline Beijing’s AI priorities. Days later at that same conference, Alibaba previewed Qwen 3.8, its own 2.4-trillion-parameter rival. The reaction in the West was a mix of awe and alarm, with observers noting that China appears to be closing what was once a comfortable American lead.

Why Kimi K3 matters

Kimi K3 is not just another model drop. With the weights public it is the largest open-weight model ever released, handing developers frontier-scale capability they can download, inspect and run themselves. That is a direct challenge to the closed, premium-priced flagships that have defined the top of the market.

The gap it closes is narrower than the headline suggests and wider than it looks.

Fourteen points on the Artificial Analysis Intelligence Index now separate K3 from Claude Opus 5.5, and 2.7 points separate K3 from the open model that overtook it in September. Two months ago the first gap was four points and the second ran the other way. The closed field moved further over the summer than the open field did, and the open field’s own leader changed hands.

The business context makes the ambition clear. Moonshot’s Kimi family reportedly passed $300 million in annualized recurring revenue by mid-June, and, according to TechCrunch’s launch report, the company was then raising fresh capital at a reported $31.5 billion valuation. Two months on, that round is reported at $50 billion with an IPO filing behind it. A well-funded lab shipping a giant open model at aggressive prices is exactly the kind of pressure that reshapes what everyone else charges. To see how fast the closed side is moving, our guide to GPT-6 Sol and Luna is a useful counterpoint.

The bottom line

Kimi K3 delivered what it promised. The weights are public, the licence is workable for almost everyone who will actually use them, and the pricing is official and low. What has not held is the crown. Artificial Analysis no longer calls K3 the strongest open model in existence, putting Xiaomi’s MiMo-V2.6-Pro and Z.ai’s GLM-5.3 above it, while Arena’s human voters still rank K3 first among open models on both boards that have seen it. It was never the best model in existence, and Moonshot said so itself.

Our recommendation, use K3 when open weights, cost or a million-token context are what you need, and when you want an open model with a real track record behind it rather than a benchmark score that is days old. Keep Claude Opus 5.5 in the loop for the hardest production coding, where it leads every independent board that has rated both. If you want the practical version of that trade-off, our Claude Opus 5.5 review covers what the closed side now offers for $4 per million input tokens.

Kimi K3 belongs in a head-to-head comparison. With Fello AI, you can run Kimi alongside Claude, GPT-6, Gemini, DeepSeek and Perplexity in one native app for Mac, iPhone and iPad. Send the same coding or research task to every frontier model at once and see which one actually wins on your work, without juggling separate accounts. When an open model this capable arrives, the fastest way to judge it is a side-by-side test. On a Mac it is a native Kimi desktop client, so K3 opens like any other app. For how K3 compares with the models inside each ChatGPT plan, see Kimi vs ChatGPT.

FAQ

Are the Kimi K3 open weights available?

Yes. Moonshot published the full Kimi K3 weights on Hugging Face on July 27, 2026, eleven days after launch. The repository holds 96 safetensors shards, roughly 1.56 TB in total, released under the custom Kimi K3 License.

Is Kimi K3 open source?

It is open-weight rather than open-source in the strict sense. The Kimi K3 License permits use, modification, distribution and commercial sale. A model-as-a-service business earning over $20 million in any 12 months needs a separate agreement, and products with over 100 million monthly users must credit Kimi K3 on screen.

How many parameters does Kimi K3 have?

Kimi K3 has 2.8 trillion total parameters in a Mixture-of-Experts design, of which 104 billion activate per token. Moonshot calls it the world’s first open 3T-class model, and its model card confirms both figures.

Is Kimi K3 better than Claude Opus 5?

Not on measured capability, and the gap widened in September. On Artificial Analysis Intelligence Index v4.3.2, Claude Opus 5 scores 50.8 and its replacement Claude Opus 5.5 scores 57.6, against 43.6 for Kimi K3, and both sit above K3 on Arena’s WebDev board. K3 wins on price, at $3 and $15 per million tokens against $5 and $25 for Opus 5 and $4 and $20 for Opus 5.5, and it is the only one of the three you can download and run yourself.

How much does Kimi K3 cost?

Official API pricing is $3.00 per million input tokens and $15.00 per million output tokens, dropping to $0.30 per million on cached input, with the cache write billed at $3.00 or $6.00 depending on how long the entry lives. Web search moved to standalone endpoints in September 2026 and bills successful calls at $0.002 to $0.003, and the model becomes available on the API platform after a top-up of at least $1.