Xiaomi MiMo is Xiaomi’s family of open-weight AI models, and on 21 September 2026 it took the top spot among open models. The flagship, MiMo-V2.6-Pro, is a 1.02-trillion-parameter mixture-of-experts model with 42 billion active parameters and a 1-million-token context window, and Xiaomi published the weights on Hugging Face under a plain MIT licence. Artificial Analysis scores it 46 on its Intelligence Index v4.3.2, ahead of GLM-5.3 at 45 and Kimi K3 at 44.

The score is not the interesting part. One Intelligence Index task costs $0.13 on MiMo-V2.6-Pro, against $2.01 on GLM-5.3 and $2.00 on Kimi K3. The new leader is roughly fifteen times cheaper per task than the two models it just passed. Below is what Xiaomi MiMo actually is, what the four V2.6 models cost, how to reach them from a Mac, and the two catches the launch coverage skipped.

The Key Takeaways

  • The flagship: MiMo-V2.6-Pro is 1.02T total and 42B active parameters, 1M context, native text, image, video and audio, under a plain MIT licence.
  • The score: 46 on the Artificial Analysis Intelligence Index v4.3.2, ranked first of 113 models in its class when we read the page on 24 September 2026.
  • The price gap: $0.13 per index task against $2.01 for GLM-5.3 and $2.00 for Kimi K3.
  • The Mac catch: MiMo Desktop runs on Windows and Apple Silicon, and Xiaomi states it is not yet available in South Korea, the UK or EU member states.
  • The free catch: the weights cost nothing, the hardware does. Xiaomi’s own SGLang recipe for Pro spans two nodes at 16-way tensor parallelism.

What Xiaomi MiMo Is, and What Shipped in V2.6

Del editor

Todos los modelos de IA en una sola app

Fello AI reúne GPT-6, Claude 5, Gemini 3.8, Grok 4.7 y más en una sola app nativa para Mac y iPhone.

¡Descárgala ahora!

MiMo is the model line built by Xiaomi’s in-house AI team. It is not a phone feature, and it is not the MIMO in your router’s spec sheet. The line started in April 2025 with a 7-billion-parameter release and has shipped a new flagship roughly every five months since. Three of those flagships arrived in 2026 alone.

The V2.6 generation went up on Hugging Face at 15:39 UTC on 21 September 2026, while Xiaomi’s own update log dates the announcement 22 September. That is why you will see both dates in coverage. Four models landed together, three of them as weights you can download.

The four models in the V2.6 family

ModelWhat it isParametersAPI price per 1M in / outBest for
MiMo-V2.6-ProFlagship reasoning model1.02T total, 42B active$0.435 / $0.87The hardest agent and coding work
MiMo-V2.6-FlashCheap high-volume workhorse309B total, 15B active$0.14 / $0.28Production traffic at scale
MiMo-V2.6-Pro-UltraSpeedLatency build of ProPro, tuned for throughput$4.35 / $8.70Real-time, latency-bound work
MiMo-V2.6-Distill-Qwen-9BResearch checkpoint on Qwen3.5-9B9BWeights onlyLocal experiments, not production

Pro and Flash are both MIT. So is the 9B distill, which Xiaomi describes on its model card as “a starting point for open research in agentic reinforcement learning”. That wording matters: it is a supervised fine-tune released for researchers, not a small MiMo you should point a product at.

Here is the launch post from Xiaomi’s own account, which is the cleanest single statement of what the company claims.

What “natively omnimodal” actually means

Most multimodal models are a text model with a vision adapter bolted on afterwards. MiMo-V2.6 puts the encoders inside the same checkpoint: a 681-million-parameter vision encoder and an audio stack of a 308-million-parameter tokenizer plus a 127-million-parameter patch encoder, all listed on the model card alongside the language backbone.

One model takes text, images, video and audio, and returns text. Artificial Analysis confirms the same input set on its own page. The practical difference is the window. All four modalities share the same 1 million tokens, so a repository, a screen recording of the bug and the call where someone described it go in together, and nothing is handed off between models.

What Xiaomi MiMo Scores, and What It Costs to Run

The number everyone quoted on launch day came from Artificial Analysis, which measures models independently rather than taking vendor figures. On version 4.3.2 of its Intelligence Index, a ten-evaluation composite, MiMo-V2.6-Pro scores 46. That ranks it first of the 113 models in its comparison class, which Artificial Analysis defines as open-weight models above 150 billion parameters, not first overall.

Put it next to the two open models it displaced and the pricing is the story.

ModelAA Intelligence Index v4.3.2Cost per index taskAPI input / output per 1M
MiMo-V2.6-Pro46$0.13$0.435 / $0.87
GLM-5.345$2.01$1.40 / $4.40
Kimi K344$2.00$3.00 / $15.00

Fifteen times cheaper, and a point ahead. That is the whole case for MiMo.

All three figures were read from Artificial Analysis on 24 September 2026, where GLM-5.3 and Kimi K3 are listed as their max variants. The index version is doing real work in that table, so quote it whenever you quote a score: v4.3.2 replaced an earlier scale, and numbers published before 21 September sit on different footing. Our ranking of every current AI model tracks the same board across the whole field.

Xiaomi’s own benchmark table, and its comparator problem

The model card carries a full evaluation table, and it is worth reading because Xiaomi does not hide the losses. Pro posts 71.9 on DeepSWE v1.1, 53.1 on AutomationBench v1.0.6, 76.9 on Toolathlon-Verified, 89.9 on Terminal Bench 2.1 and 94.0 on CyberGym. On AutomationBench and Terminal Bench 2.1 it edges past Claude Opus 5 in Xiaomi’s own testing.

There is a timing problem with that table. Xiaomi benchmarks against Claude Opus 5, GPT-5.6 Sol and Claude Fable 5. Anthropic shipped Opus 5.5 on 22 September, the day after MiMo landed, so the comparator set was one generation stale within 24 hours of publication.

Where MiMo does not lead

Claude Opus 5 stays ahead on DeepSWE v1.1, ProgramBench and Terminal Bench 4.0 in Xiaomi’s own numbers, and GPT-5.6 Sol leads on the cybersecurity evaluations ExploitBench and SEC Bench Pro. Terminal Bench 4.0 is the widest gap in the table: Xiaomi reports 34.9 for Pro against 49.0 for Opus 5. That harness is one of the ten evaluations inside the Artificial Analysis index, which is part of why the overall score lands where it does. GLM-5.3, the model it displaced, sits a single point behind on the same index.

Xiaomi concedes the ceiling itself. Its release post calls Pro “the most powerful open-source model available” and then adds, in the same sentence, that “there is still a gap when compared with the strongest closed-source models Claude Fable 5.1 and GPT-6 Astra”. A vendor marking its own limit is worth more than a vendor marking its own win.

The board humans vote on says something different again. On the LMArena text leaderboard, read the same day, Kimi K3 is the highest-placed open model at 17th with a score of 1485. MiMo-V2.6 is not on it at all. The best-placed Xiaomi entry is the previous flagship, MiMo-V2.5-Pro, down at 44th. “Top open-weight model” is true on one specific index, and a launch week is not enough voting time to know more than that.

Speed is the other place to be careful. Artificial Analysis measured Pro at 51.3 output tokens per second on 24 September, ranked 43rd of 113 and below the 66 average for its class, so the cheapest model here is also a slow one. Provider figures vary widely, and launch-day coverage quoted numbers two to three times higher, so treat any single throughput number as a snapshot of one endpoint.

Xiaomi MiMo Pricing: API Rates, Token Plan and the UltraSpeed Trade

There are two ways to pay Xiaomi and one way to pay nobody. The API is metered per token, the Token Plan is a monthly subscription aimed at coding tools, and the weights are a free download if you own the hardware.

Pay as you go

Pro runs $0.435 per million input tokens and $0.87 per million output, with a 99% cache-hit discount that drops cached input to $0.0036. Flash is $0.14 and $0.28. Those are live OpenRouter rates, and they match what Xiaomi charges directly. The company confirmed at launch that API pricing carried over unchanged from V2.5.

The cache discount is not decoration. At 99% off, a long agent session that re-reads the same repository pays almost nothing for the repeated context, which is where a 1-million-token window normally gets expensive.

The Token Plan subscription

Xiaomi sells four individual tiers at $6, $16, $50 and $100 a month, quoted in dollars and yuan on the same page, with annual packages from $63.36 to $1,056. The plans work with OpenCode, OpenClaw and Claude Code, which makes this a direct answer to subscription coding plans rather than a chat product.

One date to diary if you already subscribe: Xiaomi’s documentation says mimo-v2.5-pro and mimo-v2.5 are “officially taken offline at 10:00 on October 21, 2026 Beijing Time”. Anything pinned to a V2.5 model name stops working that morning.

UltraSpeed costs ten times as much for twenty times the speed

Pro-UltraSpeed is the variant Xiaomi markets as “up to 20x faster” at the same quality. OpenRouter, which actually serves it, describes the same model as delivering “roughly 10x the output speed”. On price the two agree: $4.35 and $8.70 per million tokens, exactly ten times the Pro rate.

So you pay a 10x premium for somewhere between a 10x and a 20x speedup, depending on which page you believe. For a batch job that is terrible value. For an interactive agent, where someone is watching a cursor blink, it is the difference between a product and a demo, and it still undercuts most closed frontier models on sticker price.

Is Xiaomi MiMo Free?

Partly, and the split is worth getting right. The weights are free: Pro, Flash and the 9B distill all sit on Hugging Face under the MIT licence, which allows commercial use and secondary training with no revenue cap and no extra permission. Xiaomi also published the technical report, more than 7,000 reinforcement-learning task environments and the training framework.

Xiaomi’s API is not free, and neither is serving the weights yourself. The company’s own deployment page for MiMo-V2.6-Pro gives an SGLang recipe spanning two nodes at 16-way tensor parallelism, and a vLLM recipe running 8-way on one node. Either way you are renting a rack, not buying a workstation. That is the honest reading of “free” for a trillion-parameter model: the licence costs nothing and the GPUs cost a fortune. Flash targets a single 8-GPU node, and the 9B distill runs on one card.

How to Use Xiaomi MiMo on a Mac or Anywhere Else

There are four routes in, and which ones are open to you depends on where you live.

MiMo Desktop, and where you cannot get it

Xiaomi shipped a desktop app of its own, version 0.1.0, on 1 September 2026. It runs on Windows and on Apple Silicon Macs. It also does the things a 2026 agent app does: browser control, cross-session task handover, document and slide generation, and a router that picks between the flagship and the cheap model by task complexity.

Then there is the line on Xiaomi’s own download page: “Not yet available in South Korea, the UK or EU member states.” A second restriction sits further down, where desktop-level computer control is marked “Global only”. If you are reading this from Berlin, Dublin or Manchester, the flagship Mac experience for Xiaomi MiMo is closed to you today, and the launch coverage did not mention it.

The API, OpenRouter and coding tools

The API has no such restriction and is the practical route for most people outside China. All three V2.6 models are live on OpenRouter with the 1-million-token context intact, which also means they drop into anything that already speaks the OpenRouter or OpenAI-compatible protocol. Xiaomi’s Token Plan plugs the same models into Claude Code and OpenCode directly.

Running the weights yourself

Self-hosting is the answer if the point of choosing an open model was keeping data off someone else’s servers. Xiaomi publishes SGLang and vLLM recipes on the model card, and both got day-zero support. Flash is the realistic target at 309 billion total parameters on a single 8-GPU node. Pro is heavier again, and no Mac runs either of them. If local inference on Apple hardware is what you are after, our guide to open-source models worth running on an M5 Mac covers what actually fits in unified memory.

The shortcut if you just want the models on a Mac

Most people asking how to use Xiaomi MiMo on a Mac do not want a GPU cluster. They want a native app that holds several strong models behind one window. Fello AI does that on macOS, iPhone and iPad, with GPT, Claude, Gemini, Perplexity, Grok, Kimi, DeepSeek, GLM and Qwen side by side, plus web search, PDF chat, and Word, Excel and PowerPoint output. MiMo is not in that roster today. Treat it as the answer to the general version of the question, and use Xiaomi’s API or Desktop when MiMo itself is the requirement.

Flash Is the Model Most People Should Actually Run

The headline belongs to Pro. The downloads belong to Flash. When we checked Hugging Face on 24 September, Flash had 13,243 downloads against 4,070 for Pro, a three-to-one gap three days after launch.

That tracks with what Flash is: 309 billion total parameters with 15 billion active, the same 1-million-token context, the same four input modalities, at a fifth of Pro’s price. It also came within a point or two of Pro on most of Xiaomi’s own agent benchmarks, posting 52.3 on AutomationBench against Pro’s 53.1 and 87.6 on Terminal Bench 2.1 against 89.9.

Tim Dettmers, the Carnegie Mellon professor who wrote the bitsandbytes quantization library, ran it and put the comparison in the class where it belongs.

His throughput figure is for Flash on his own setup, which is why it sits so far above the 51.3 tokens per second Artificial Analysis measured for Pro. Different model, different harness. It is a useful reminder that one published speed number describes one endpoint and nothing else.

From Hunter Alpha to the Top of the Open Board

This is the second time a Xiaomi model has surprised the field. In March 2026 an unnamed model called Hunter Alpha appeared on OpenRouter and topped the daily charts for several days. It processed over a trillion tokens, and had a good part of the internet convinced it was an unreleased DeepSeek. It was MiMo-V2-Pro, and we covered the mystery model that turned out to be Xiaomi’s when the identity came out.

That model shipped as a proprietary API product. The line has opened up since: MiMo-V2.5-Pro arrived under MIT in April 2026, and V2.6 continues it with the weights, the training environments and the report all published together. The 2026 pattern across the best open-source AI models is Chinese labs shipping weights first and arguing about licences second, and DeepSeek V4 set much of that template.

Luo Fuli, a former DeepSeek researcher who now leads the MiMo team, framed the V2.6 run on X as a deliberate bet: “In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL.” Xiaomi puts the bill at $2.62 million for Pro and $850,000 for Flash, with the two models completing 30 reinforcement-learning steps each and roughly 750,000 trajectories between them in under six days.

Should You Use Xiaomi MiMo?

If you buy tokens by the million and your workload is agentic, yes, and you should be testing Flash this week rather than Pro. A fifteenfold cost advantage at the same rough intelligence tier is not a benchmark curiosity, it is a budget line, and the 99% cache discount widens it further on long sessions.

If you need the fastest possible answer, or you are in the EU or the UK and wanted the desktop app, look elsewhere for now. And if you were told a trillion-parameter model is now free, the licence is, the two-node cluster is not. Xiaomi MiMo earned the top of the open board on price more convincingly than on capability, and price is the harder thing for the rest of the field to answer.

Frequently Asked Questions

Is Xiaomi MiMo free to use?

The weights are free. MiMo-V2.6-Pro, Flash and the 9B distill are all published under the MIT licence, which permits commercial use and retraining without extra permission. Xiaomi’s API is metered, at $0.435 per million input tokens for Pro and $0.14 for Flash, and self-hosting Pro means a multi-GPU server at minimum.

Can I download MiMo Desktop on a Mac?

Yes, on Apple Silicon, with one large exception. Xiaomi’s download page states the app is not yet available in South Korea, the UK or EU member states, and desktop-level computer control is marked as available in the global build only. Readers in those regions have to use the API instead.

Is Xiaomi MiMo better than DeepSeek or Kimi K3?

On the Artificial Analysis Intelligence Index v4.3.2, MiMo-V2.6-Pro scores 46 against GLM-5.3 at 45 and Kimi K3 at 44, and costs $0.13 per index task against roughly $2 for both. On the LMArena text leaderboard, where humans vote, Kimi K3 is the highest-placed open model at 17th and MiMo-V2.6 has no votes yet.

What hardware do I need to run MiMo-V2.6 locally?

More than a Mac. Xiaomi’s published SGLang recipe for Pro uses 16-way tensor parallelism across two nodes, and its vLLM recipe runs 8-way on a single node. Flash, at 309 billion total parameters, targets a single 8-GPU node, and the MiMo-V2.6-Distill-Qwen-9B checkpoint runs on one GPU, though Xiaomi releases it as a research starting point rather than a finished model.

When do the older MiMo models stop working?

Xiaomi’s Token Plan documentation says mimo-v2.5-pro and mimo-v2.5 are taken offline at 10:00 on 21 October 2026, Beijing Time. Anything pinned to those model names needs to move to the V2.6 series before that morning.