The best open source AI models changed hands again in September. Xiaomi published the weights for MiMo-V2.6-Pro on September 21, 2026 under a plain MIT license, and Artificial Analysis now scores it 46.3 on its Intelligence Index v4.3.2, the highest of any open-weight model and level with Grok 4.7. The top closed model, Claude Opus 5.5, sits at 57.6. Kimi K3 is still the largest open model and the open model Arena’s voters rank highest, GLM 5.3 Flash is the best MIT-licensed pick for most people, and Gemma 4 still runs on a laptop. Our full DeepSeek vs Claude comparison puts one open lab against a closed one side by side.

This guide ranks the open source AI models worth your time in 2026, what each one is genuinely best at, and how to actually run them without a rack of GPUs. We ranked the picks on independent benchmark data, license terms, and real accessibility, because most “best open source AI” lists are written by companies selling you the hardware to run them. We sell an app, not GPUs, so this ranking has no horse in that race. For a closer look at how an open model holds up against the paid flagships, see GLM head-to-head with the closed-source flagships.

If the “open source” label itself feels fuzzy, our explainer on what open source AI actually means breaks down open weights, licensing, and why “open” does not always mean what you think.

The Key Takeaways

  • MiMo-V2.6-Pro is the highest-scoring open model on the Artificial Analysis Intelligence Index at 46.3, a 1.02-trillion-parameter MIT-licensed model that costs about $0.13 per index task. Kimi K3 (43.6) is still the largest open model at 2.78 trillion parameters and the top open model on Arena’s human-vote boards, where MiMo has not been rated yet.
  • GLM 5.3 Flash is the right pick for most people: 41.8 on the index, a 1M-token context, native image and video input, and a clean MIT license on a 321B model you can actually rent hardware for. The older GLM 5.2 scores 33.7 and Artificial Analysis now lists it as deprecated.
  • DeepSeek V4.1 Flash (39.5, MIT) replaced V4-Pro as DeepSeek’s strongest open model on September 10, and it is the fastest model in this ranking at about 233 tokens per second.
  • Most “open source” models are actually open-weight, the weights are free but the training data stays private. True open-source AI is rarer than the label suggests, and several of the strongest models here ship under custom licenses rather than MIT or Apache 2.0.
  • You do not need a GPU farm for most of this list. Gemma 4 12B runs on a modern laptop, Qwen 3.8 27B fits a well-equipped Mac, and apps route the biggest models to your Mac or iPhone with zero local hardware.

What Counts as an Open Source AI Model

Od vydavatele

Každý AI model v jedné aplikaci

Fello AI přináší GPT-6, Claude 5.5, Gemini 3.8, Grok 4.7 a další v jedné nativní aplikaci pro Mac a iPhone.

Stáhnout hned!

An open source AI model publishes its weights, and ideally its code and training data, so anyone can download, run, modify, and redistribute it. That is the promise. The reality is more nuanced, and it matters before you pick one.

Most models everyone calls “open source” are really open-weight. You get the trained weights for free, but the training data and full recipe stay locked away. The Open Source Initiative’s official definition requires far more transparency than that, which is why almost no frontier model qualifies as fully open. We flag the distinction throughout, because a “Modified MIT” or custom community license can carry commercial strings that a true MIT or Apache 2.0 license does not.

The practical takeaway is simple. Open weights mean you can run the model privately, fine-tune it, and avoid per-token API fees. They do not always mean unrestricted commercial use, so the license column in our table below is just as important as the benchmark scores.

Best Open Source AI Models 2026 at a Glance

Here is the full ranking. We picked these ten on independent benchmark performance, recency, license freedom, and how realistic they are to actually use. The index column is the Artificial Analysis Intelligence Index v4.3.2, read on September 25, 2026. Scores marked “est.” are the board’s own estimates rather than full runs, and scores from earlier versions of the index are not comparable with these, so an older “57” for Kimi K3 is the same model on a different ruler. Chinese labs hold the top five places. Thinking Machines’ Inkling, the strongest US-built family, gets its own section below.

Model Best for Params (total / active) Context License AA Index v4.3.2
MiMo-V2.6-Pro Highest measured score 1.02T / 42B 1M MIT 46.3
Kimi K3 Largest open model, top on Arena 2.78T / 104B 1M Kimi K3 (custom) 43.6
GLM 5.3 Flash Best for most people 321B / 18B 1M MIT 41.8
DeepSeek V4.1 Flash Speed and value 552B / 16B 1M MIT 39.5
Qwen 3.8 27B Best model you can run on a Mac 27.8B dense 262K Apache 2.0 33.7
Kimi K2.7 Code Agentic coding on less hardware 1T / 32B 256K Modified MIT 25.8
Muse Glimmer Meta’s open model for local agents 30B dense 131K Apache 2.0 17.5
Gemma 4 31B Lightweight + local 31B dense 256K Apache 2.0 19.0 (est.)
Nex-N2-Pro Fully permissive agentic 397B / 17B 262K Apache 2.0 28.2 (est.)
MiniMax M3 Efficiency on light hardware 427B / 23B 1M MiniMax Community 29.2

The Best Open Source AI Models, Ranked

Here is each pick in detail, in order, with what it is genuinely best at, the benchmarks that earned its spot, and the license you actually get. We start with the strongest all-rounders and work down to the most efficient and the most local.

1. MiMo-V2.6-Pro, the Highest-Scoring Open Model

MiMo-V2.6-Pro from Xiaomi took the top of the open field on September 21, 2026, the day its weights went up on Hugging Face. Xiaomi’s model card describes a sparse Mixture-of-Experts design with 1.02 trillion total and 42 billion active parameters, a 1-million-token context, and native text, image, video and audio input in one model. The license is plain MIT, with no revenue threshold and no attribution clause, which neither of the two open models closest to it on the index can say.

The independent numbers are what put it first. Artificial Analysis scores it 46.3 on Intelligence Index v4.3.2, ahead of GLM-5.3 at 44.8 and Kimi K3 at 43.6, and level with xAI’s closed Grok 4.7 at 46.4. It is also the cheapest model near the top of that board, at about $0.13 per index task against $2.00 for Kimi K3. The caveat is that nobody has voted on it yet: MiMo-V2.6-Pro has no entry on Arena’s text or WebDev boards, so the human-preference picture is still K3’s. Our Xiaomi MiMo explainer covers the whole family, its API prices and the smaller V2.6 Flash.

2. Kimi K3, the Largest Open Model and Arena’s Open Favourite

Kimi K3 from Moonshot AI became a download on July 27, 2026, and it is still the biggest open model you can get. Hugging Face lists 2.78 trillion parameters, a Mixture-of-Experts design that activates 104 billion per token, handles text, images and video natively, and carries a 1-million-token context window. The nearest rival on size is Alibaba’s Qwen 3.8 flagship at 2.45 trillion.

The independent boards now split on it. Artificial Analysis scores K3 at 43.6 on Intelligence Index v4.3.2, third among open-weight models behind MiMo-V2.6-Pro and GLM-5.3. Arena’s voters still put it first in the open field: it is ninth overall on WebDev at 1,659.6, with only closed models above it, and 17th on the text board, again the highest open entry. Moonshot’s own launch table adds 88.3 on Terminal-Bench 2.1 and 81.2 on FrontierSWE, and those two are vendor-run numbers rather than independent ones.

Two things to know before you plan around it. The weights ship under Moonshot’s own Kimi K3 License, which is permissive for almost everyone. It only bites if you resell model access above $20 million in revenue, which needs a separate agreement, or pass 100 million monthly users, which needs prominent “Kimi K3” credit. The other catch is hardware, because 2.78 trillion parameters arrive as 96 weight shards, so this is a server model and not a laptop one. Our Kimi K3 breakdown has the full benchmark table.

3. GLM 5.3 Flash, the Best Open Source Model for Most People

GLM 5.3 Flash is the open model most teams should actually run. Released by Z.ai on August 26, 2026, it is a 321-billion-parameter Mixture-of-Experts model that activates about 18 billion per token, reads a 1-million-token context, takes images and video as well as text, and ships under a clean MIT license. It scores 41.8 on the Artificial Analysis index at about $0.25 per task, so it is within five points of the top of the open field on well under half the parameters of anything above it.

It replaces GLM 5.2, which held this slot until September. GLM 5.2 now scores 33.7, Artificial Analysis marks it deprecated, and Z.ai charges the same $1.40 / $4.40 per million tokens for it as for the full GLM 5.3, which scores 44.8. GLM 5.3 is the stronger model and its weights are on Hugging Face, but they ship under a custom license rather than MIT, which is why the Flash model gets the recommendation. The rest of Zhipu’s GLM family is covered in our explainer, and GLM 5.5 is rumoured for later in 2026.

4. DeepSeek V4.1 Flash, Best for Speed and Value

DeepSeek V4.1 Flash replaced V4-Flash on September 10, 2026, and it is now the strongest open model DeepSeek makes. The model card describes a multimodal Mixture-of-Experts model with 552 billion backbone parameters that activates only 8 billion per token while reading input and 16 billion while writing, with native image input, a 1-million-token context and an MIT license. Meituan’s open-sourced LongCat-2.0 is another 1M-token option, built specifically for agentic coding.

It scores 39.5 on the Artificial Analysis index at about $0.27 per task, and at roughly 233 tokens per second it is the fastest model in this ranking. The bigger V4-Pro (1.6 trillion total, 49 billion active) is now the weaker of the two at 36.0 and has no image input, so its April launch scores, 80.6 on SWE-Bench Verified and 93.5 on LiveCodeBench, describe a model DeepSeek has already overtaken itself. Read the full DeepSeek V4 breakdown for both models and the current peak and off-peak API prices.

5. Qwen 3.8 27B, the Best Open Model You Can Run on a Mac

Alibaba’s Qwen 3.8 27B is the strongest model on this list that fits on a single well-equipped machine. It is a 27.8-billion-parameter dense model under a clean Apache 2.0 license, released on Hugging Face in August 2026, with native image and video understanding and a 262K-token context that Alibaba says extends to 1 million. It scores 33.7 on the Artificial Analysis index, exactly level with GLM 5.2, a model 27 times its size. The larger Qwen 3.8 flagship, 2.45 trillion parameters, also has public weights and scores 39.9, but it ships under Alibaba’s own license, as does the 180B Qwen3.8-Flash-Next (39.8), the second-highest open model on Arena’s WebDev board.

The earlier Qwen 3.6 models are still downloadable, and the 35B-A3B is still the lighter option for modest hardware, but the 27B has moved on. Alibaba’s strongest model overall, Qwen3.8 Max, scores 45.4 and is API-only and closed-weight, so it sits outside this open-source ranking. For one open model that does a bit of everything on hardware you own, Qwen 3.8 27B is the safest bet. If you run it yourself, the Prague lab BottleCap AI publishes ThinkingCap fine-tunes that cut its reasoning tokens by 37% to 46% for under a point of accuracy.

6. Kimi K2.7 Code, Best for Agentic Coding on Less Hardware

K3 is the flagship, but Kimi K2.7 Code is the Kimi most teams can realistically deploy, because it is under two-fifths of K3’s size. It is a 1-trillion-parameter model activating 32 billion per token, with a 256K-token context window under a Modified MIT license, released June 12, 2026 and purpose-built for agents that write, run, and debug across many steps. You can see what Moonshot’s hosted models cost in our Kimi pricing guide.

One honest caveat. Most of what Moonshot published, including a 62.0 on its own Kimi Code Bench v2 and 81.1 on MCP Mark Verified, comes from in-house benchmarks. The first independent reading is less flattering: Artificial Analysis scores it 25.8 on its general index, well below the models above it, though that index is not a coding-only test. Treat it as a specialist rather than an all-rounder. Our Kimi K2.7 Code review has the details and the missing-data flags.

7. Muse Glimmer, Meta’s Open Model for Local Agents

Meta’s open-weight release this year is Muse Glimmer, published on August 10, 2026 under a fully permissive Apache 2.0 license. It is a 30B dense model (about 29.8 billion parameters, including its vision encoder) distilled from Meta’s closed Muse Spark, with a 131K context window and image input, and Meta built it for agents that use tools and recover from their own mistakes rather than for chat. There is no Llama 5. Meta’s last Llama generation was Llama 4 in April 2025, and its 2026 models ship under the Muse name, with the flagship Muse Spark kept closed. Meta has promised open weights for Muse Spark too, a promise it repeated when Muse Spark 1.3 shipped on 2 September 2026, still without a date.

What makes Glimmer notable is that Meta benchmarked it on Macs itself. Its 4-bit build shrinks the language model to under 20 GB, the smallest official quant is aimed at 24 GB machines, and on Meta’s own numbers it runs at 37.8 tokens per second on an M4 Max and 50.2 on an M5 Max with its DFlash speculative decoding switched on. Ollama ships an Apple Silicon build you can pull with ollama run muse-glimmer:30b-mlx. The independent picture is modest, 17.5 on the Artificial Analysis index, and Meta’s own table is mixed: it beats Gemma 4 31B clearly on MCP Atlas (75.5 against 54.2) and SWE-Bench Pro (51.2 against 36.9) but loses to Qwen 3.6 27B on OSWorld-Verified and Terminal-Bench 2.1. It is a strong local agent model, not a straight upgrade over everything in its size class.

8. Gemma 4, Best Lightweight Model for Local Use

If you want to run AI on your own machine with the least fuss, Gemma 4 from Google DeepMind is the answer. The family spans tiny edge models (E2B, E4B) and a 12B model up to a 31B dense flagship and a 26B MoE variant, all under Apache 2.0, released April 2, 2026.

Despite the small footprint it punches hard on the boards people vote on. Gemma 4 31B sits at 1,451 on Arena’s text board, above MiniMax M3, Inkling and Qwen 3.8 27B, even though Artificial Analysis estimates it at only 19.0 on its tougher index. Google’s own figures are 85.2% on MMLU-Pro, 89.2% on AIME 2026, and 80.0% on LiveCodeBench v6, while the 12B tier runs comfortably on a modern laptop. It is the best entry point if you are new to local AI, and Google’s official Gemma 4 model card lists exact specs per size.

9. Nex-N2-Pro, Best Fully Permissive Agentic Model

Nex-N2-Pro from Nex AGI earns its spot on license freedom plus genuine capability. It is a 397-billion-parameter MoE model (17B active) built on Qwen 3.5’s largest open model, with a 262K context, shipped under the fully permissive Apache 2.0 license and tuned specifically for agentic work like tool-calling and long multi-step tasks.

Nex AGI reports 80.8 on SWE-Bench Verified and 75.3 on Terminal-Bench 2.1, and Artificial Analysis estimates it at 28.2 on its index, in the same band as MiniMax M3 while carrying zero commercial restrictions. See our Nex-N2-Pro analysis for how it stacks up against the top closed models.

10. MiniMax M3, Best for Efficiency

MiniMax M3 rounds out the list for anyone watching hardware cost. Released June 1, 2026, it runs 427 billion total parameters with only 23 billion active, and its MiniMax Sparse Attention design decodes roughly 15x faster at full context than M2. MiniMax reports 59.0% on SWE-Bench Pro, Artificial Analysis scores it 29.2 at about $0.51 per task, and it serves at around 153 tokens per second, with a full 1-million-token context and native multimodality.

The license changed with this release. M3 ships under the MiniMax Community License, free to self-host but asking larger commercial users to arrange a separate agreement, a step back from the Modified MIT terms of earlier versions. It is still a smart pick when you want fast, frontier-adjacent coding on a leaner hardware budget. Our MiniMax model review covers the lineage and the benchmark caveats.

Inkling, the Strongest US Open-Weights Model

Inkling is the first model from Thinking Machines, the lab founded by former OpenAI CTO Mira Murati, and it landed on July 15, 2026. It is a Mixture-of-Experts transformer with 975 billion total parameters and 41 billion active per token, a context window of up to 1 million tokens, and native reasoning over text, images and audio. The weights ship under a plain Apache 2.0 license, which puts it in the most permissive tier on this page alongside Qwen 3.8 27B, Gemma 4, Muse Glimmer and Nex-N2-Pro.

The headline is national rather than global. On Artificial Analysis Intelligence Index v4.3.2, Inkling scores 25.0, ahead of NVIDIA’s Nemotron 3 Ultra at 22.9, Muse Glimmer at 17.5 and gpt-oss-120b at 11.6, which keeps the family at the top of the US open field. It is also more than 20 points behind MiMo-V2.6-Pro, so the story is that the best American open model leads the American field, not the world. On its own scorecard Inkling posts 77.6% on SWE-bench Verified and 73.5% on MMMU Pro.

Where it earns attention is speed and cost. Artificial Analysis measures Inkling at about 167 tokens per second and $0.61 per index task, cheaper per task than Kimi K3 or GLM-5.3 and roughly three to five times faster than either, though MiMo-V2.6-Pro and GLM 5.3 Flash now score far higher for less.

The more interesting release is the small one. Inkling-Small shipped its full weights on July 30, 2026 at 276 billion total parameters and 12 billion active, under a third of its parent, and Artificial Analysis estimates it at 27.8, above the flagship. It also edges the larger model on Humanity’s Last Exam (33% against 32%) and GPQA Diamond (89% against 87%) and serves at about 205 tokens per second. If you want a US-built model on hardware you can realistically rent, Inkling-Small is the more practical half of this pair.

We have left Inkling out of the numbered ranking above because on capability alone it would sit near the bottom of it, while on license freedom it matches the best on this page. Treat it as the default pick when you specifically want a US-built model on a fully permissive license. Tencent’s 770B Hy4 Preview arrived under Apache 2.0 on August 28, 2026 and is listed here rather than ranked, since no independent board has rated it yet.

How We Picked These Open Source AI Models

We ranked on four things, in order. Independent benchmark performance, led by the Artificial Analysis Intelligence Index and Arena’s human-vote boards, with lab-reported tests (SWE-Bench Verified, GPQA Diamond, MMLU-Pro, LiveCodeBench) as supporting evidence. Recency, every model here shipped or updated in 2026. License freedom, with a clear preference for MIT and Apache 2.0 over custom community licenses. And accessibility, because a model you cannot realistically run is not useful to you. Our guide to AI benchmarks explains what each of those tests actually measures.

We also weighted independent verification. Where scores come only from the lab that built the model, as with parts of Moonshot’s K3 table and most of Nex AGI’s, we say so, and where Artificial Analysis only estimates a score we mark it. Self-reported benchmarks are a real problem in open AI right now, and you should discount them until third parties confirm the numbers.

You Do Not Need a GPU Farm to Use These

The biggest myth about open source AI is that running it requires server-grade hardware. It depends entirely on the model. The trillion-parameter giants like Kimi K3 and MiMo-V2.6-Pro do need multi-GPU setups for full-precision inference, but plenty of strong open models run on consumer gear.

Gemma 4 12B, the smaller Qwen tiers and Meta’s Muse Glimmer run on a modern laptop or Mac using tools like Ollama or LM Studio, and you can grab the weights straight from Hugging Face. And if you want the biggest models without owning any hardware at all, an app like Fello AI routes leading models straight to your Mac or iPhone, no local install required. Running a model yourself also sidesteps provider outages entirely, a real concern given Claude’s repeated 2026 downtime. Three other systems are worth knowing about even though they are not ranked here. ByteDance’s Seed 2.1 Pro is an unreleased preview that reached #8 on Code Arena: Frontend in June. Sakana AI’s Fugu orchestrates a pool of other models instead of competing as a single set of weights, and its Fugu-Ultra v1.1 update widened that lead on coding and terminal benchmarks. Quantized versions also shrink memory needs dramatically, often at minimal quality cost.

If coding is your main use case, our roundup of the best free AI for coding compares these open models head-to-head on developer tasks specifically.

Conclusion

MiMo-V2.6-Pro is now the highest-scoring open source AI model you can download, and it carries a plain MIT license. Kimi K3 is the largest, and the open model human voters rank highest. For most people GLM 5.3 Flash is the better choice, with the strongest blend of reasoning, license freedom, and a 1M-token context on hardware you can actually rent. If you want speed and value, go DeepSeek V4.1 Flash. If you want to run AI on your own Mac, start with Qwen 3.8 27B or, on lighter hardware, Gemma 4. If you want a US-built model on a clean Apache 2.0 license, Inkling-Small is the best value in that corner right now. And if you would rather skip the setup entirely, the same frontier open models are available through the best AI models on apps like Fello AI.

The gap between the best open model and the best closed one is 11.3 points on Artificial Analysis Intelligence Index v4.3.2, 46.3 for MiMo-V2.6-Pro against 57.6 for Claude Opus 5.5. That is a real gap at the very top, but the open leaders now sit level with closed models like Grok 4.7 for a fraction of the price, so “free” is often the smarter choice. Pick by your use case, check the license, and you will not miss the subscription. Check the release terms too, because GLM 5.3 and its staged weights show that an open line can ship its best model under stricter terms than the one before it.

FAQ

What is the best open source AI model right now?

MiMo-V2.6-Pro is the highest-scoring open model as of September 25, 2026, at 46.3 on the Artificial Analysis Intelligence Index v4.3.2, with MIT-licensed weights published on September 21. Kimi K3 is the largest open model and the top open entry on Arena’s human-vote boards. GLM 5.3 Flash is the better choice for most people, with an MIT license and a fraction of the hardware bill. Among US-built open models, Thinking Machines’ Inkling family leads.

What is the difference between open-source and open-weight AI?

Open-weight means the trained weights are free to download and run, but the training data and full recipe stay private. True open-source AI, by the Open Source Initiative definition, also releases the data and code. Most “open source” models are technically open-weight.

Can I run open source AI models without a GPU?

Yes. Small models like Gemma 4 12B run on a modern laptop, Qwen 3.8 27B and Muse Glimmer run on a well-equipped Mac, and apps like Fello AI let you use the largest open models from a Mac or iPhone with no local hardware. Only the biggest trillion-parameter models need multi-GPU setups.

Are open source AI models free for commercial use?

Usually, but check the license. MIT and Apache 2.0 models (MiMo-V2.6-Pro, GLM 5.3 Flash, DeepSeek V4.1 Flash, Qwen 3.8 27B, Gemma 4, Muse Glimmer, Nex-N2-Pro, Inkling and Inkling-Small) are fully permissive. Custom licenses like the Kimi K3 License, the GLM 5.3 license, Alibaba’s Qwen 3.8 flagship license or the MiniMax Community License add conditions, such as revenue thresholds, usage caps or attribution requirements.

Are open source models as good as ChatGPT or Claude?

Close, but not level at the top. MiMo-V2.6-Pro scores 46.3 on the Artificial Analysis Intelligence Index v4.3.2 against 57.6 for Claude Opus 5.5, the top closed model on that board, and GPT-6.1 Sol sits well above the open leader at 51.8. For many everyday coding, math, and long-context tasks the difference is small, but the absolute frontier still belongs to the closed labs.