Best AI Models in 2026
Rankings, comparisons, and deep dives, updated monthly as new models ship.
Claude Fable 5.1 arrived on September 1 and went straight to the top of Artificial Analysis's Intelligence Index at 66, the highest score the house has measured, taking its Agentic Index at 61 with it. Claude Opus 5 keeps the two things most readers are actually choosing on: the coding crown, won on both of Arena's vote-based boards, and the price, $5 / $25 per 1M tokens against Fable 5.1's $10 / $50. No Arena board has votes for 5.1 yet, so nothing this page decides on human preference has moved. Grok 4.6 was August's real surprise: a post-training refresh of Grok 4.5 that jumped five points to 61, tying GPT-5.6 Sol while charging $2 / $6, and finishing long agentic jobs in roughly half the turns Opus 5 needs.
The other move is at the cheap end, and for the first time this year it went the wrong way. DeepSeek raised both V4 models roughly fourfold on August 16 and split the rate card into peak and off-peak tiers, so DeepSeek V4-Flash 0731 now costs $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak instead of the flat $0.14 / $0.28 it launched on, which costs it the price-performance pick. That pick moves to GLM 5.3 Flash, which landed on August 26 with MIT weights on day one, an Intelligence Index of 57 and a list price of $0.15 / $0.50, halved to $0.075 / $0.25 until September 9, and which now undercuts DeepSeek at every hour of the day. Google had already halved its own Flash rate: Gemini 3.7 Flash shipped on August 13 at an introductory $0.75 / $3.75, with Gemini 3.6 Flash moved onto the same rate, and Gemini 3.8 Flash followed on September 2 at exactly that price, scoring 59 on Artificial Analysis's Intelligence Index against 56 for 3.7 Flash. All three prices double on January 1, 2027. The open field then closed the month with three releases in three days: GLM 5.3 finally published its weights on August 28, though under a bespoke Z.ai licence rather than the plain MIT its smaller sibling got, Alibaba opened Qwen3.8-Flash-Next on August 26 as the first public look at the Qwen4 architecture, and Tencent open-sourced Hy4 Preview on August 28, 770 billion parameters under Apache 2.0. Below are the category winners for September 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks, and pricing are updated within 48 hours of any major model launch.
Want the top AI models without juggling separate subscriptions? Fello AI brings the leading models together in one native app for Mac, iPhone and iPad.
Download Fello AISep 2 Google Gemini 3.8 Flash, three points on the Intelligence Index at the same $0.75 / $3.75 New
Sep 1 Anthropic Claude Fable 5.1 and Mythos 5.1, a new #1 on Artificial Analysis at 66 and a 75% cut to cache reads New
Aug 28 Z.ai GLM 5.3 weights are out after the safety hold, but the licence is not MIT New
Aug 28 Tencent Tencent Hy4 Preview, 770B open weights under Apache 2.0 and the largest permissive model yet New
Aug 26 Z.ai GLM 5.3 Flash, Ox Alpha unmasked, and the first MIT weights Z.ai has shipped since GLM-5.2 New
Aug 26 Alibaba Qwen3.8-Flash-Next, open weights and the first public look at the Qwen4 architecture New
Aug 21 OpenAI GPT-5.6 Sol cut by more than 20%, now under Claude Opus 5 on both sides New
Aug 16 DeepSeek DeepSeek raises API prices roughly fourfold and splits the rate card into peak and off-peak tiers New
Aug 14 Z.ai GLM 5.3, a large agentic jump from post-training alone, shipped without the open weights New
Aug 13 Google Gemini 3.7 Flash, coding scores jump and the price halves to $0.75 / $3.75, until January New
Aug 12 SpaceXAI Grok 4.6, post-training refresh takes Intelligence Index 56 to 61 at an unchanged $2 / $6 New
Aug 5 Muse Spark 1.2 and Muse Code, Intelligence Index 51 to 57 at an unchanged $1.25 / $4.25, plus a terminal coding agent New
Aug 3 Alibaba Qwen 3.8 (Qwen3.8-Max), 2.4-trillion-parameter multimodal flagship now generally available at $2 / $6 New
Jul 31 DeepSeek DeepSeek V4-Flash 0731, same price, same size, Intelligence Index 40 to 52
Jul 31 MiniMax MiniMax H3, 2K video with native stereo audio from $0.13 a second
Late Jul Thinking Machines Inkling Small, a quarter the size of Inkling at nearly the same score
Jul 30 OpenAI GPT-5.6 price cut, Luna 80% cheaper, Terra 20% cheaper, Sol unchanged
Jul 24 Anthropic Claude Opus 5, new #1 on Artificial Analysis, tops both the Intelligence Index (63) and the Agentic Index (55.3)
Jul 24 Sakana AI Fugu-Ultra v1.1, orchestration-engine refresh with vendor-reported gains of up to 7.9 points over v1.0 at the same price
Jul 21 Google Gemini 3.6 Flash, cheaper output, faster, same Intelligence Index
Jul 16 Moonshot AI Kimi K3, 2.8T params, the largest open-weight model ever released (weights shipped July 27)
Jul 9 OpenAI GPT-5.6 Sol, Terra, and Luna, next-gen family live across ChatGPT, Codex, and the API
Pending Google Gemini 3.5 Pro, still unreleased, months behind schedule per Bloomberg Delayed
Best AI for Writing
The best AI for writing is Claude Fable 5, the only model in the top three of all three independent writing boards, with Claude Sonnet 5 as the free-tier value pick and GPT-5.5 as the alternative for fact-anchored business writing. Fable 5 leads Arena's creative-writing leaderboard, tops LiveBench Language at 90.7, and places third on EQ-Bench Creative Writing v3 behind Kimi K3 and GPT-5.6 Sol. No other model is top-three on more than one of them. This is a change from last month, when Claude Sonnet 5 held this slot on GDPval-AA; the preference boards do not support that placing, so Sonnet 5 is now the value pick rather than the quality leader. Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. If your writing is a work deliverable rather than prose, the GDPval-AA v2 professional-deliverables board is now led by Claude Fable 5.1 at 1853, ahead of Claude Opus 5 at 1824 and Fable 5 at 1723. Fable 5.1 landed on September 1 as the higher-capability model at the same $10 / $50, and it beats Fable 5 on every board that has measured both, including Humanity's Last Exam at 59.1% against 55.5%. We have not moved the crown to it, because this pick rests on human preference and no Arena board has rated it yet. Translation is a separate question with a separate winner, which we work through in our guide to the best AI for translation.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Fable 5 | Best writing overall | #1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #3 EQ-Bench | Priciest option here | $10 / $50 |
| Claude Fable 5.1 | Highest capability at the same price | #1 Humanity's Last Exam (59.1%), #1 GDPval-AA v2 (1853), #1 Intelligence Index (66) | No Arena votes yet, so unrated on the preference boards | $10 / $50 |
| Kimi K3 | Creative fiction and voice | #1 EQ-Bench Creative Writing (2377), 234 Elo clear of second | Only #10 on Arena creative writing | $3 / $15 |
| Claude Sonnet 5 | Free everyday writing | Free and default on claude.ai, 1M context | #53 Arena creative writing; 75.0 LiveBench Language | $2 / $10, now permanent |
| Claude Opus 5 | Professional deliverables on a budget | #2 GDPval-AA v2 at 1824, ahead of Fable 5 (1723) | Behind Fable 5 on Arena text, behind Fable 5.1 on GDPval-AA v2 | $5 / $25 |
| GPT-5.5 | Fact-anchored business writing | Documented factual-reliability gains over GPT-5.4 | Reasoning tiers now marked deprecated | $5 / $30 |
| Gemini 3.8 Flash | Bulk drafts at scale | Intelligence Index 59, 304.6 tok/s output, 1M context | Agent-tuned rather than prose-tuned; intro price doubles January 1, 2027 | $0.75 / $3.75 |
The writing crown stays with Claude Fable 5, but the case for it narrowed on September 1. Anthropic shipped Claude Fable 5.1, and it takes Humanity's Last Exam at 59.1% against Fable 5's 55.5%, the GDPval-AA v2 board at 1853, and the Intelligence Index at 66. What it does not have is a single Arena vote, and this crown rests on Arena's text and creative-writing boards, where Fable 5 is still #1 at 1508.6 on the August 1 cutoff. So Fable 5.1 goes into the table as the higher-capability option at the same $10 / $50, and we re-test the crown the moment Arena rates it. Two numbers we quoted last month have also been re-fitted: AA-Omniscience now reads 43 for both Fable models rather than 40, and Artificial Analysis measures Fable 5 on Humanity's Last Exam at 55.5% rather than 53.3%.
Best AI for Chat & Daily Assistant
The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena's text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $4 / $20 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| GPT-5.6 | Everyday chat, ChatGPT's default | The assistant most people can open; Terra matches GPT-5.5 at ~half cost | #14 on Arena text; scheming flagged by METR | Free / $20/mo Plus; API $0.20 / $1.20 to $4 / $20 |
| Claude Fable 5 | Highest-rated conversation | #1 on Arena text overall (1508.6) and 6 of 7 subcategories | No free tier; usage-credit access on Pro | $10 / $50 API |
| Claude Opus 5 | Thoughtful, nuanced answers | #1 Artificial Analysis Intelligence Index (63) and Agentic Index (55.3) | #6 on Arena text, behind Claude Fable 5 | $20/mo Pro, $5 / $25 API |
| GPT-5.5 Instant | Hallucination-sensitive daily work | 52.5% fewer hallucinated claims vs 5.3 Instant | Reasoning tiers now marked deprecated | $20/mo Plus; API $5 / $30 |
| Gemini 3.6 Flash | Fast, free, multimodal | Still the model the free Gemini tier serves, 1M context, #12 on Arena text | Weaker on hardest reasoning; 3.8 Flash outscores it at the same price | Free / $0.75 / $3.75 API |
| Fello AI | All the top models, one app | ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad | Routed via app, not direct | $9.99/mo |
GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open.
Best AI for Images
The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena's text-to-image board at 1385 Elo and its image-editing board at 1463, and Artificial Analysis puts it first on its own image arena at 1339.4. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro. The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta's Muse Image is #3 on Arena's text-to-image board and #2 on image editing, the strongest showing any Meta image model has managed. Google's Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| ChatGPT Images 2.0 | Images with readable text | #1 text-to-image (1385) and #1 image editing (1463) | Less photoreal than the Gemini image line | Included in ChatGPT Plus |
| Reve 2.1 | Layout, typography, native 4K | #2 text-to-image at 1302, layout-preserving editing | Smaller ecosystem | Free / from $7.99/mo |
| Muse Image | Image editing, Meta ecosystem | #3 text-to-image, #2 image editing (1407) | New, thin tooling around it | Meta AI app |
| Nano Banana 2 (Gemini 3.1 Flash Image) | Photoreal portraits and products | Outranks Nano Banana Pro on both boards | Weaker on text in image | Gemini app / AI Studio |
| Seedream 5.0 Pro | Multilingual text + region editing | 10+ languages incl. Arabic RTL, lasso and layer editing | No independent benchmarks; copyright cloud | BytePlus / Magnific |
| Midjourney v8 | Stylized art, illustration | Aesthetic baseline most artists prefer | Weaker on text in image | $10-$120/mo |
| Grok Imagine | NSFW / Spicy Mode | Most permissive guardrails | Smaller model behind it | $30/mo SuperGrok |
The image crown is unchanged and GPT Image 2 still leads all three boards we track, on Arena text-to-image (1385), Arena image editing (1463) and Artificial Analysis's image arena (1339.4). The one addition is Microsoft's MAI-Image-2.5, which sits third on Artificial Analysis's image board at 1269.7 and third on Arena's image-editing board, so third place now depends on which house you read.
Best AI for Video
The best AI for video generation is Gemini Omni Flash, which leads Arena's text-to-video board at 1527 Elo, a full 45 points clear of second place, and also tops Artificial Analysis's video arena. It is #1 on both houses, which no other video model manages. Pricing runs $1.50 in and $17.50 per 1M video output tokens, which works out at roughly $0.10 per second of finished video, and it supports conversational editing. The one real limit is length: Omni Flash generates 10-second clips. If you need longer takes, Veo 3.1 remains the right tool inside the Gemini app, AI Studio and Vertex AI, with native audio and 1080p output. This replaces Veo 3.1 at the top of the category; Veo 3.1 is a good model but not the leading one, with its best variant at #6 on Arena's text-to-video board. The contender to watch is MiniMax H3 (July 31), which generates 2K clips of 4 to 15 seconds with native stereo audio and is billed per second from $0.13 at 2K, already #2 on Artificial Analysis's video board at 1241.5. We are not moving the crown on that: the margin is 3.3 Elo on the single board that has not reproduced between parses, and H3 still has no votes on Arena's text-to-video board.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Gemini Omni Flash | Best AI video overall | #1 on both video boards (1527 Arena), conversational editing | Caps at 10-second generations | ~$0.10/sec; Gemini app / AI Studio |
| MiniMax H3 | 2K clips with native audio | 4-15s at 2K, native stereo audio; #2 on AA video (1241.5) | No Arena votes yet; the open weights exclude the 2K upscaler and need an application in the US, EU, UK and South Korea | $0.13/sec at 2K |
| Dreamina Seedance 2.0 | Closest challenger | #2 text-to-video (1482) and #1 on image-to-video with audio (1198) | ByteDance ecosystem, limited Western access | Dreamina / BytePlus |
| Muse Video | Meta ecosystem video | #3 text-to-video at 1459 | Newest of the group, thin tooling | Meta AI app |
| Veo 3.1 | Longer production clips | Native audio, 1080p, strong physics consistency | #6 on Arena video, not the quality leader | Google AI Pro / Ultra |
| Kling 3.0 / 3.0 Turbo | Fast iteration at lower cost | Native 4K, 60fps, 15-second clips; Turbo shipped June 17 | Outside the top 16 on Arena text-to-video | From $10/mo |
| Luma Ray 3 | Photoreal scenes | Strong realism for landscapes | Smaller community | Free / from $9.99/mo |
Gemini Omni Flash keeps the video crown, and it is still #1 on both boards, at 1527.5 on Arena and 1244.8 on Artificial Analysis. MiniMax H3 (July 31) enters as the contender at #2 on Artificial Analysis's video board (1241.5), 3.3 Elo behind, and it beats Omni Flash on specification with 2K output, clips up to 15 seconds and native stereo audio. We are holding the crown because that margin is 3.3 Elo on a single board, and because H3 still has no votes on Arena's text-to-video board. MiniMax did open the weights on August 3, but only H3-Base and the VAEs and encoder; the 2K upscaler that the specification advantage rests on stayed API-only.
Best AI for Coding
The best AI for coding is Claude Opus 5, and it wins on the two boards where developers vote on the finished result rather than a script measuring a harness. It is #1 on Arena's WebDev board at 1,702.9 and #1 on image-to-WebDev at 1,668.6, on vote cutoffs of August 1 and July 31. Price is the second half of the argument: Opus 5 runs $5 / $25 per 1M tokens against Claude Fable 5's $10 / $50, and Artificial Analysis measures it at $2.34 per index task against Fable 5's $3.14. Anthropic's own docs tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for workloads that need the highest available capability. Claude Fable 5 is the runner-up and stays the pick for the hardest long-horizon work, at #4 on WebDev (1,630.7) and #2 on image-to-WebDev (1,625.7); its September 1 successor, Claude Fable 5.1, is the stronger model on the harness benchmarks at the same price, reaching 55.8% on Terminal-Bench 4.0 against Opus 5's 52.3% and Fable 5's 42.0% on Anthropic's own figures, but Arena has no votes for it yet. The contender is Kimi K3, which takes #2 on WebDev at 1,675.5 and #3 on Arena's Agent board, the strongest open-weight coder on the boards, though self-hosting it means 1.56 TB of weights. On Artificial Analysis's Coding Index it is not a Claude sweep: GPT-5.6 Sol (xhigh) leads at 78.3 with Opus 5 (max) at 78.0, close enough to call a tie. The cheapest serious contender is now GLM 5.3 Flash at $0.15 / $0.50 under MIT, after DeepSeek's August 16 rate rise took V4-Flash 0731 to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak, and the best value at the frontier is Grok 4.6, which posts 88.4% on Terminal-Bench v2.1 and $0.84 per index task at $2 / $6.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Opus 5 | Best coding overall | #1 Arena WebDev (1,702.9) and #1 image-to-WebDev (1,668.6); $2.34 per index task | Artificial Analysis Coding Index has GPT-5.6 Sol a shade ahead | $5 / $25 |
| Claude Fable 5 | Hardest long-horizon agentic work | #2 Arena image-to-WebDev (1,625.7), #4 WebDev; 1M context | Priciest; Artificial Analysis Coding Index puts it 7th | $10 / $50 |
| Claude Fable 5.1 | Highest-capability agentic coding | #1 Intelligence Index (66) and #1 Agentic Index (61); 55.8% Terminal-Bench 4.0 (Anthropic) | No Arena votes yet; twice the price of Opus 5 | $10 / $50 |
| Kimi K3 | Web app building and agents | #2 Arena WebDev (1,675.5), #3 Arena Agent, highest open Intelligence Index (60) | 1.56 TB to self-host | $3 / $15 |
| GPT-5.6 Sol | OpenAI flagship, agentic coding | #1 Artificial Analysis Coding Index at 78.3 (xhigh) | Absent from the official Terminal-Bench board; eval-gaming flagged by METR | $4 / $20 |
| Muse Spark 1.2 | Cheap agentic coding | Intelligence Index 57 at $1.25 / $4.25; 80% on Terminal-Bench 2.1 (Artificial Analysis) | Second to Opus 5 on all three of Meta's own coding charts (82.9% vs 86.7%); Scale SEAL has rated only 1.1 | $1.25 / $4.25 |
| Grok 4.6 | Cheap value coder | 88.4% on Terminal-Bench v2.1 and Intelligence Index 61, at $0.84 per index task | Prompts of 200K tokens and up re-bill the whole request at double; higher hallucination rate | $2 / $6 |
| Gemini 3.8 Flash | Agent coding at scale | Intelligence Index 59, 73.7% DeepSWE v1.1, 89.4% Terminal-bench 2.1, 304.6 tok/s | 19.1% on Terminal-bench 4.0 against Opus 5's 51.8%; intro price doubles January 1, 2027 | $0.75 / $3.75 |
| GLM 5.3 Flash | Best open-weight coder you can host | Intelligence Index 57; DeepSWE v1.1 63.4 and Terminal Bench 2.1 84.3 (vendor figures); MIT, 320B/18B active | Kimi K3 outscores it; the coding figures are vendor-reported and it has no Arena votes yet | Open weights (MIT) |
The coding crown moved to Claude Opus 5. Arena finished rating it, and it came first on both coding boards, WebDev at 1,702.9 and image-to-WebDev at 1,668.6, with Claude Fable 5 fourth and second. Those are vote-based boards with published cutoffs, so the crown now rests on human comparisons rather than on any one harness, and Opus 5 does it at half Fable 5's price. Fable 5 becomes the runner-up for the hardest long-horizon work, and Kimi K3 stays the contender at #2 on WebDev. Meta's Muse Spark 1.2 arrived on August 5 and does not change that, because on Meta's own charts it comes second to Opus 5 on every coding benchmark it published. Grok 4.6 (August 12) does not take the crown either, but it is the one to watch on cost: it reaches 88.4% on Terminal-Bench v2.1 and finishes long agentic jobs in roughly half the turns Opus 5 needs, at $2 / $6 against $5 / $25. Two more coding-focused launches landed right after it. Gemini 3.7 Flash (August 13) replaces 3.6 Flash in this table on the strength of a jump from 49.0% to 65.3% on DeepSWE v1.1 at half the old price, and GLM 5.3 (August 14) posts the largest agentic gain of the month, Terminal-Bench 3.0 from 4.6 to 28.3, but shipped with no downloadable weights. Z.ai closed the month by opening a different model instead: GLM 5.3 Flash (August 26) arrived under MIT on day one at Intelligence Index 57, and it takes the open-weight coding slot from GLM-5.2. The last days of the month then produced three more open releases: GLM 5.3's own weights on August 28 under a bespoke licence rather than MIT, Alibaba's Qwen3.8-Flash-Next on August 26 as a preview of the Qwen4 architecture, and Tencent's 770B Hy4 Preview on August 28 under Apache 2.0. None of them has an independent coding rating yet, so none of them moves this table. Anthropic then opened September with Claude Fable 5.1 on the 1st, which tops Artificial Analysis's Intelligence and Agentic indexes and posts 55.8% on Terminal-Bench 4.0 against Opus 5's 52.3%, but Arena has not rated it, so the crown stays where the votes are. Google closed the gap at the cheap end on September 2 with Gemini 3.8 Flash, which takes the at-scale row from 3.7 Flash at the same price: DeepSWE v1.1 goes from 65.3% to 73.7% and Terminal-bench 2.1 from 85.8% to 89.4%. It does not touch the crown, because Opus 5 still wins Terminal-bench 4.0 by 51.8% to 19.1%.
Best AI for Creativity
The best AI for unfiltered, on-trend creative work is Grok 4.6, and we want to be exact about why. This pick is about the product, not the prose quality. Grok carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, and the boards were blunt about its predecessor: Grok 4.5 sat #33 on EQ-Bench Creative Writing and #41 on Arena's creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. Neither creative-writing board has rated Grok 4.6 yet, and since SpaceXAI aimed this release at coding and agentic work rather than prose, we are not assuming it moved. If you are picking on output quality alone, Claude Fable 5 wins outright at #1 on Arena creative writing, and Kimi K3 wins EQ-Bench. Choose Grok 4.6 for what it will let you make, not for how well it writes.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Grok 4.6 | Unfiltered, opinionated, on-trend | Fewest content restrictions, native real-time X grounding | Creative-writing boards have not rated it; Grok 4.5 sat #33 EQ-Bench, #41 Arena | $30/mo SuperGrok |
| Claude Fable 5 | Highest-quality creative prose | #1 Arena creative writing, #1 LiveBench Language | Cautious guardrails on edgy briefs | $10 / $50 API |
| Kimi K3 | Fiction and distinctive voice | #1 EQ-Bench Creative Writing at 2377 | Only #10 on Arena creative writing | $3 / $15 |
| Claude Opus 5 | Long-form structured creativity | Holds long threads and self-edits; Intelligence Index 63, second only to Claude Fable 5.1 | Most cautious of the group | $20/mo Pro; $5 / $25 API |
| Gemini 3.1 Pro | Multimodal creative | Strong text, image and video chain | Quotas inside the Gemini app | Free / $2.00-$4.00 API in |
| Grok Imagine (Spicy Mode) | NSFW / adult creative | Most permissive image generation | Niche use case | $30/mo SuperGrok |
Grok keeps this pick, now on Grok 4.6, which replaced Grok 4.5 as SpaceXAI's flagship on August 12. Nothing about the product's permissiveness or its live X access changed, which is what the pick rests on. The model underneath is meaningfully stronger, up five points to Intelligence Index 61, but that gain landed in coding and agentic work rather than prose, and no creative-writing board has rated 4.6 yet. Claude Fable 5 remains the model to use when you want the better writing.
Best AI for Accuracy & Research
The best AI for accuracy and research is Gemini 3.1 Pro. Its strongest result is on ARC Prize's ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you want when the answer has to be current rather than merely plausible. It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity's Last Exam, and tops Scale SEAL's HLE board at 46.44. We have dropped the previous ARC-AGI-2 framing: its 77.1% is still correct, but the board has moved and that score now places it around 14th. Two honest caveats. On grounded search specifically, Arena's search leaderboard is led by Anthropic, not Google, with Gemini 3.1 Pro grounding at #7. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience, a board Claude Fable 5.1 now shares the top of.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| Gemini 3.1 Pro | Cheap, reliable factual work | 98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol) | ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search | $2.00-$4.00 / $12.00-$18.00 (tiered) |
| Claude Fable 5 | Grounded search | #1 and #3 on Arena's search leaderboard, ahead of Google | No single cheap tier | $10 / $50 (Fable 5) |
| GPT-5.6 Sol | Novel reasoning | #1 ARC-AGI-2 at 92.5%, against a 100% human panel | Scheming flagged by METR | $4 / $20 |
| Claude Opus 5 | Hardest unseen problems | #1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize) | Artificial Analysis measures a 50% hallucination rate | $5 / $25 |
| Qwen 3.7 Max | Frontier accuracy at value pricing | 92.4 GPQA Diamond, 200 free requests/day | API-only, no chat front-end | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Opus 4.6 | Honesty under pressure | #1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5 | Superseded as a flagship | Legacy Anthropic model |
Gemini 3.1 Pro keeps the accuracy crown, but its headline number is no longer a lead. Its GPQA Diamond score re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly, so the crown rests on the ARC-AGI-1 result at $0.52 a task plus Search grounding rather than on GPQA. This is the next crown we re-test, and the boards that measure how often a model is simply wrong moved on September 1. Claude Fable 5.1 now holds the highest factual accuracy Artificial Analysis has measured, at 67% against Claude Fable 5's 65%, and tops Humanity's Last Exam at 59.1% against 55.5%. The AA-Omniscience index itself reads 43 for both models, up from the 40 we quoted for Fable 5 last month, because 5.1 buys its extra accuracy with a higher attempt rate on questions it cannot answer.
Best AI for Problem Solving
The best AI for hard problem solving is GPT-5.6 Sol, and it is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 92.5%, the closest any model has come to the 100% human panel. Three separate houses put it first on the reasoning tasks that matter, which is more agreement than any other category on this page produces. OpenAI has still not published Sol's FrontierMath score, so the verified OpenAI mark remains GPT-5.5 Pro's 39.6% on FrontierMath Tier 4, and we will slot Sol's number in the moment it goes public. Qwen 3.7 Max is the value alternative for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, at a fraction of the cost of ChatGPT Pro, and it now includes 200 free model requests per day. Claude Opus 5 is the alternative for long agentic reasoning chains, at 59 on Artificial Analysis's Agentic Index, second to Claude Fable 5.1's 61, and #1 on ARC Prize's ARC-AGI-3 at 30%, roughly 3.75x the next-best model.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| GPT-5.6 Sol | Hardest math, science and reasoning | #1 LiveBench Mathematics (96.2), #1 LiveBench Reasoning (91.7), #1 ARC-AGI-2 (92.5%) | FrontierMath still unpublished; scheming flagged by METR | $100/mo ChatGPT Pro; API $4 / $20 |
| Claude Opus 5 | Long agentic reasoning chains | #2 Agentic Index (59), #1 ARC-AGI-3 (30%) | Second on LiveBench Reasoning; 50% hallucination rate | $5 / $25 |
| GPT-5.5 Pro | Verified FrontierMath leader | 39.6% FrontierMath Tier 4 | Superseded by Sol as flagship | $100/mo ChatGPT Pro |
| Qwen 3.7 Max | Competition math on a budget | 97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day | API-only | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Fable 5 | Math inside a coding workflow | #1 on Arena's math subcategory (1543), 96.0 LiveBench Mathematics | Priciest option here | $10 / $50 |
| GLM 5.3 Flash | Open-weight problem solving | Intelligence Index 57 under MIT, four points above GLM-5.2 at less than half the size; 320B/18B active, 1M context | Needs a multi-GPU server, not a workstation; GLM 5.3 scores higher on Z.ai's own figures and is downloadable since August 28, but at more than twice the size and under a bespoke licence | Open weights (MIT) |
GPT-5.6 Sol keeps this crown, and it picked up a second argument. Sol now also leads Artificial Analysis's Terminal-Bench v2.1 at 89.5% and its Coding Index at 78.3, and it matches Gemini 3.1 Pro on GPQA Diamond at 94.1%. Claude Opus 5 stays the agentic-reasoning alternative on the strength of ARC-AGI-3. On August 1 OpenAI named its next model family Astra and published ten new mathematical results from an internal version, though Astra is unreleased and has no date, price or availability. Anthropic's Claude Fable 5.1 arrived on September 1 and took the Agentic Index off Opus 5, 61 against 59, which changes the alternative's supporting number rather than the crown.
Best AI Agent
The best AI agent right now is Gemini Spark for 24/7 cloud-resident work and Claude Cowork for desktop-resident work, with ChatGPT Codex as the alternative for coding agents and OpenAI Operator-class browser agents as the alternative for web tasks. AI agents are the fastest-moving category of 2026: each top vendor now ships an agent product, and the practical choice is between agents that live in the cloud (run while your laptop is closed) and agents that live on your desktop (drive your apps directly). Gemini Spark launched at Google I/O on May 19, 2026 and is the first 24/7 cloud agent. Claude Cowork launched in general availability on April 9, 2026 and runs as a desktop agent that drives your local apps. ChatGPT Codex Mobile (May 14) is the pick for coding-agent work. Read the full Gemini Spark vs Claude Cowork comparison.
| Agent | Best For | Where It Runs | Strength | Price |
|---|---|---|---|---|
| Gemini Spark | 24/7 cloud tasks, Workspace workflows | Google Cloud VM (always-on) | First true 24/7 agent, deep Workspace integration | $99.99/mo Google AI Ultra |
| Claude Cowork | Desktop, app-driving, design + code | Your Mac/Windows desktop | Drives local apps, sees your screen | $20/mo Claude Pro |
| ChatGPT Codex Mobile | Coding agent on phone | OpenAI cloud + iOS/Android | Approve diffs and redirect work from phone | Included in ChatGPT plans |
| Grok Agentic (Grok 4.6) | Real-time research, X scraping | SpaceXAI cloud | Native X integration; GDPval-AA v2 Elo 1755 | $30/mo SuperGrok |
| OpenAI Operator-class | Browser tasks, web forms | OpenAI cloud + your browser | Web automation | ChatGPT Pro |
No new consumer agents shipped, and Meta's Muse Code (August 5) is a terminal developer tool rather than a consumer one, so the Gemini Spark (cloud) versus Claude Cowork (desktop) choice still drives most agent decisions for individual users. The model layer underneath them moved again: Claude Fable 5.1 (September 1) now leads Artificial Analysis's Agentic Index at 61, ahead of Claude Opus 5 at 59, and tops the GDPval-AA v2 board at 1853 against Opus 5's 1824. Opus 5 is still the one most teams should build on at $5 / $25, half Fable 5.1's price, and it tops Arena's Agent board, with Kimi K3 in third. For teams building their own agents, Meta's Muse Spark 1.2 (August 5) is a cheap agent-native option at $1.25 / $4.25 whose GDPval-AA v2 agentic Elo jumped from 1371 to 1631, though Scale SEAL has rated only the older 1.1, which leads its MCP Atlas tool-use board at 88.1. GLM 5.3 Flash (August 26, MIT) is now the open-weight agent model we would host, at Intelligence Index 57 and AutomationBench 48.8 on Z.ai's own chart, where Claude Opus 4.8 scores 41.0, though it has no Arena votes yet; GLM-5.2 keeps the independent record, at #5 on Arena's Agent board. Grok 4.6 (August 12) is the new cost story in this category rather than a new agent product: it posts a GDPval-AA v2 Elo of 1755, which the September entries pushed down to fifth among distinct models, and Artificial Analysis measures it finishing long-horizon jobs in roughly 53 turns and 0.5 billion input tokens against about 103 turns and 2.0 billion for Opus 5, at $0.84 per index task.
The leading models like ChatGPT, Claude and Gemini, together on Mac, iPhone and iPad.
Free to start, 4.7★ across 27,000+ reviews.
Download Fello AIBest AI for Students
The best AI for students is GPT-5.5 Free inside ChatGPT for general coursework and Gemini 3.6 Flash Free inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Opus 5 as the alternative for essay editing. Most students don't need to pay: the free ChatGPT tier now defaults to GPT-5.6 (with GPT-5.5 still available), Gemini 3.6 Flash is in the free Gemini app and AI Studio, Claude Sonnet 5 is the new free Claude default, and DeepSeek V4 is free on DeepSeek's chat site. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's new flagship (its FrontierMath score is not yet published, with GPT-5.5 Pro's verified 39.6% on FrontierMath Tier 4 the current mark), though both are paid-only; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).
| Task | Best Model | Why | Free? | Alternative |
|---|---|---|---|---|
| Essays & coursework | GPT-5.5 | Free in ChatGPT, improved factual reliability vs 5.4 | Yes | Claude Sonnet 5 (free Claude) |
| STEM problem-solving | GPT-5.6 Sol / Qwen 3.7 Max | New STEM flagship (5.5 Pro: 39.6% FrontierMath) / 97.1 HMMT 2026 Feb | Pro paid / Qwen API paid | Gemini 3.6 Flash (free) |
| Research & accuracy | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, native Google Search grounding | Yes (Gemini app) | Claude Opus 5 |
| Writing editing | Claude Sonnet 5 | Free and default on claude.ai; Claude Fable 5 is the quality leader | Yes (Claude free) | Claude Fable 5 |
| Multimodal study (PDFs, slides, images) | Gemini 3.6 Flash | 1M context, free in Gemini app | Yes | NotebookLM (Google) |
Best AI for Work & Professionals
The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark on Google AI Ultra at $99.99/month is the only true 24/7 cloud agent.
| Use Case | Best Model | Key Stat | Price | Alternative |
|---|---|---|---|---|
| Daily knowledge work | GPT-5.6 | ChatGPT's default since July 9; the assistant most people can open | $20/mo ChatGPT Plus | Claude Opus 5 |
| Coding (proprietary) | Claude Opus 5 | #1 on Arena WebDev (1,702.9) and image-to-WebDev (1,668.6); Anthropic's recommended default | $20/mo Claude Pro | Claude Fable 5 |
| Coding (cost-effective) | Qwen 3.7 Max | 80.4 SWE-Verified, 1M context | $1.25 / $3.75 promo; $2.50 / $7.50 list | DeepSeek V4-Flash 0731 |
| Research & briefings | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, Google grounding | Google AI Pro / Ultra | Claude Opus 5 |
| Hard math, physics, finance modelling | GPT-5.6 Sol | OpenAI's new STEM flagship (5.5 Pro verified at 39.6% FrontierMath) | $100/mo ChatGPT Pro | Qwen 3.7 Max |
| Always-on agent workflows | Gemini Spark | First 24/7 cloud agent | $99.99/mo Google AI Ultra | Claude Cowork |
| Live news, X-context creative | Grok 4.6 | Intelligence Index 61 + native X grounding | $30/mo SuperGrok | Gemini 3.1 Pro |
| All-in-one consolidation | Fello AI | ChatGPT + Claude + Gemini + Grok + DeepSeek | $9.99/mo | Pay each vendor separately |
AI Model Pricing in September 2026
From $0 free tiers to $199.99/month Google AI Ultra. The most consequential price on this table is Claude Opus 5 at $5 / $25, because it is the second-highest scorer on Artificial Analysis's Intelligence Index at half the cost of the two Fable models around it. Claude Fable 5.1, which took the top of that index on September 1, holds the same $10 / $50 list price as Fable 5 but cuts cache reads 75% to $0.25, which is where the saving on long agentic runs actually sits rather than on the sticker. Meta's Muse Spark 1.2 lists at $1.25 / $4.25, with a Contributor tier at $0.10 / $0.20 for anyone willing to let Meta train on their prompts, and Grok 4.6 lists at $2 / $6, with the caveat that a prompt of 200,000 tokens or more re-bills the entire request at $4 / $12. The price-performance pick moved in August, and not because anything got cheaper: DeepSeek raised both V4 models roughly fourfold on August 16, taking V4-Flash 0731 from a flat $0.14 / $0.28 to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak. The pick is now GLM 5.3 Flash at $0.15 / $0.50 list, halved until September 9, which scores 57 on the Intelligence Index and undercuts even DeepSeek's off-peak rate on both sides at every hour of the day. At the frontier itself the value pick is Grok 4.6, scoring 61 for $0.84 per index task where Opus 5 costs $2.34 and GPT-5.6 Sol charges $4 / $20 for the same score after its August 21 cut. Among closed models GPT-5.6 Luna is the cheapest at $0.20 / $1.20 after OpenAI's July 30 cut. For a deeper breakdown see our full AI Pricing Comparison Guide.
| Model | Input (per 1M) | Output (per 1M) | Context | Free Access? |
|---|---|---|---|---|
| GPT-5.5 | $5.00 | $30.00 | 1M (400K in Codex) | ChatGPT Free; API paid |
| GPT-5.5 Pro | $30.00 | $180.00 | 1M | ChatGPT Pro from $100/mo ($200 higher-usage tier) |
| GPT-5.6 Sol | $4.00 | $20.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Terra | $2.00 | $12.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Luna | $0.20 | $1.20 | Not published | Live in ChatGPT, Codex & API (July 9) |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Claude Pro/Max default; API paid |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Legacy model at Anthropic; Pro/Max/API |
| Claude Fable 5.1 | $10.00 | $50.00 | 1M | Cache reads $0.25; in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits. Mythos 5.1 bills identically, by invitation only |
| Claude Fable 5 | $10.00 | $50.00 | 1M | Permanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits |
| Claude Sonnet 5 | $2.00 | $10.00 | 1M | Claude Free & Pro default; API paid |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | API paid (superseded by Sonnet 5) |
| Gemini 3.1 Pro | $2.00 (≤200K) / $4.00 (>200K) | $12.00 (≤200K) / $18.00 (>200K) | 1M | Limited Gemini app; API paid |
| Gemini 3.8 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | AI Studio, Antigravity, Gemini Enterprise + paid API; in the Gemini app reported for AI Pro/Ultra |
| Gemini 3.7 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | AI Studio, Android Studio, Antigravity + paid API; in the Gemini app via Spark only (AI Pro/Ultra, excludes EEA/UK/CH/Nigeria) |
| Gemini 3.6 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | Free Gemini app default; AI Studio; free API tier + paid API |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | AI Studio; free API tier + paid API |
| Qwen 3.7 Max | $1.25 promo / $2.50 list | $3.75 promo / $7.50 list | 1M | 200 free requests/day; API paid beyond that |
| MiniMax M3 | $0.30 (50% off $0.60) | $1.20 (≤512K) | 1M | Open weights; hosting costs apply |
| LongCat-2.0 | Provider-dependent | Provider-dependent | 1M | Open weights (MIT); hosting costs apply |
| NVIDIA Nemotron 3 Ultra | Provider-dependent | Provider-dependent | 1M | Open weights (OpenMDW); hosting costs apply |
| Qwen 3.5 (open-weight) | Self-host / Together | Self-host / Together | 1M | Open weights; hosting costs apply |
| Nex-N2-Pro | Self-host / providers | Self-host / providers | 1M | Open weights (Apache 2.0); hosting costs apply |
| Rio 3.5 Open 397B | Self-host / providers | Self-host / providers | 1M | Open weights (MIT); hosting costs apply |
| Grok 4.3 | $1.25 | $2.50 | 1M | Free consumer plan; API paid |
| Grok 4.6 | $2.00 (<200K) / $4.00 (≥200K) | $6.00 (<200K) / $12.00 (≥200K) | 500K | SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel, Cloudflare |
| Muse Spark 1.2 | $1.25 ($0.10 Contributor tier) | $4.25 ($0.20 Contributor tier) | 1M | Paid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount |
| Kimi K3 | $3.00 ($0.30 cache-hit) | $15.00 | 1M | Free basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License) |
| Gemini Omni Flash (video) | $1.50 | $17.50 (video output) | 10-second clips | Gemini app / Flow; AI Studio + API |
| DeepSeek V4-Pro | $0.66 off-peak / $1.32 peak ($0.022 cache-hit off-peak) | $1.98 off-peak / $3.96 peak | 1M | DeepSeek Chat free; API paid, peak/off-peak tiers since August 16 |
| DeepSeek V4-Flash 0731 | $0.22 off-peak / $0.44 peak ($0.007 cache-hit off-peak) | $0.66 off-peak / $1.32 peak | 1M | DeepSeek Chat free; API paid, peak/off-peak tiers since August 16 |
| Kimi K2.7 Code | Provider-dependent | Provider-dependent | 256K | Open weights; hosting costs apply |
| GLM-5.2 | Provider-dependent | Provider-dependent | 1M | Open weights; hosting costs apply |
| GLM 5.3 | $1.40 ($0.26 cache-hit) | $4.40 | 1M | Z.ai API, GLM Coding Plan and ZCode; weights on Hugging Face since August 28 under the bespoke GLM 5.3 License |
| GLM 5.3 Flash | $0.15 ($0.03 cache-hit); $0.075 promo | $0.50; $0.25 promo | 1M | Z.ai API, GLM Coding Plan, OpenRouter; MIT weights on Hugging Face; promo ends September 9 |
| Tencent Hy4 Preview | $0.834 ($0.042 cache-hit) | $2.501 | 1M+ | Open weights (Apache 2.0); Tencent Cloud TokenHub, OpenRouter, WorkBuddy, CodeBuddy |
| Qwen3.8-Flash-Next | Provider-dependent | Provider-dependent | 262K (to 1M) | Open weights (Qwen Community License 1.0); hosting costs apply |
| ERNIE 5.1 | China-region pricing | China-region pricing | 256K | Baidu free tier |
| Gemini Spark (agent) | Not API-priced | Not API-priced | 1M (Gemini base) | Google AI Ultra $99.99 or $199.99/mo |
| Fello AI (aggregator) | Routed via app | Routed via app | Model-dependent | $9.99/mo, free tier available |
The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. Grok 4.6 works the same way but bites harder: at 200,000 prompt tokens and above, SpaceXAI re-bills the entire request at the higher rate rather than only the overage, so crossing the line by a thousand tokens doubles the cost of the whole call. DeepSeek is the other rate card that needs reading twice: since August 16 both V4 models bill at peak rates from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays, and at exactly half that in every other hour, so the same job can cost twice as much depending on when you run it. If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.
Best Open-Weight Models in September 2026
The best open-weight model in September 2026 is Kimi K3, and it took the lead the moment Moonshot published the weights on July 27, 2026. It holds the highest Intelligence Index of any open model at 60 on Artificial Analysis, seven points clear of GLM-5.2, and on Arena it is still the highest-placed open entry, at #2 on WebDev (1,675.5) behind only Claude Opus 5 and #3 on the Agent board. One catch decides which of the two you should actually use. The K3 download is 96 safetensors shards and about 1.56 TB, which needs a multi-node GPU cluster rather than a workstation, and it ships under a custom Kimi K3 License rather than MIT. So the practical recommendation for teams running their own weights is now GLM 5.3 Flash (Z.ai, MIT), released on August 26 at Intelligence Index 57, four points above GLM-5.2 and at less than half the size, 320B total with 18B active. It is the first natively multimodal model in the GLM-5 line, it takes a 1M-token context, and the weights were on Hugging Face under MIT the day it launched. GLM-5.2 (MIT) is the fallback, at Intelligence Index 53, and it still holds the independent record the newer model has not had time to earn, #6 on Arena WebDev (1,587.1) and #5 on the Agent board. Third place changed hands on the last day of July: DeepSeek V4-Flash 0731 was re-post-trained and jumped from Intelligence Index 40 to 52 on the current v4.1.1 index under MIT, though its API price has since gone up roughly fourfold, to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak. Three releases then closed the month, and none of them has an independent rating yet. GLM 5.3 finally published its weights on August 28, on exactly the date its placeholder had carried, but under a bespoke GLM 5.3 License rather than MIT, with a clause requiring any model-as-a-service operator above $10 billion in revenue to pass a Z.ai security review first; note also that GLM 5.3 Flash is not a trimmed GLM 5.3 but a separate model on a newly trained base, which is why it could ship under plain MIT while the flagship could not. Tencent open-sourced Hy4 Preview on August 28 at 770B total and 49B active under Apache 2.0, the largest permissively licensed model published so far, on the strength of an internal blind evaluation rather than any independent board. And Alibaba published Qwen3.8-Flash-Next on August 26, a 125B mixture-of-experts with only 6B active that previews the Qwen4 architecture, under the Qwen Community License 1.0 rather than Apache. All three are listed below rather than ranked.
| Model | Best For | Key Benchmark | Context / License | Where To Run |
|---|---|---|---|---|
| Kimi K3 | Highest-scoring open model, agentic and web work | II 60, highest of any open model; #2 Arena WebDev, #3 Arena Agent; 2.8T/104B active | 1M / Kimi K3 License | Hugging Face (96 shards, 1.56 TB), Moonshot API, providers |
| GLM 5.3 Flash | Best open model most teams can actually run | II 57 (v4.1.1), on the intelligence-versus-cost Pareto frontier at about $0.09 an index task; 320B/18B active, natively multimodal | 1M / MIT | Hugging Face, Z.ai API ($0.15/$0.50), OpenRouter |
| GLM-5.2 | Fallback with an independent track record | II 53; #6 Arena WebDev (1,587.1), #5 Arena Agent; 744B/40B active | 1M / MIT | Z.ai, Hugging Face, OpenRouter |
| GLM 5.3 | Open since August 28, but on a bespoke licence | Same 744B base as GLM-5.2, post-trained only, published checkpoint counted at 753B; CyberGym 84.5% and Terminal-Bench 3.0 28.3 (vendor figures) | 1M / GLM 5.3 License (MaaS above $10B revenue needs a Z.ai security review) | Hugging Face, Z.ai API ($1.40/$4.40), GLM Coding Plan, ZCode |
| DeepSeek V4-Flash 0731 | Best open value if you host it yourself | II 52 (v4.1.1), up from 40 on July 31; 284B/13B active | 1M / MIT | DeepSeek API ($0.22/$0.66 off-peak, $0.44/$1.32 peak since August 16), local |
| Tencent Hy4 Preview | Largest permissively licensed model yet | 770B/49B active; 2.99/4.00 on Tencent's own 203-task expert panel vs Kimi K3 2.94 and GLM 5.3 2.92 (vendor); no independent rating yet | 1M+ / Apache 2.0 | Hugging Face (BF16 and FP8), Tencent Cloud TokenHub, OpenRouter |
| Qwen3.8-Flash-Next | Preview of the Qwen4 architecture | 125B total plus a 51B N-gram embedding, 6B active; 91.7 GPQA Diamond and 62.5 SWE-Bench Pro (vendor); no independent rating yet | 262K to 1M / Qwen Community License 1.0 | Hugging Face, providers, self-host |
| LongCat-2.0 | Frontier open coder trained on Chinese chips | II 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active | 1M / MIT | Hugging Face, GitHub, OpenRouter |
| MiniMax M3 | Cheap frontier-class multimodal | II 44 (v4.1), 59% SWE-Bench Pro, multimodal | 1M / license TBD | Hugging Face, API $0.30/1M (50% off) |
| Nex-N2-Pro | Strongest open coding score | II 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B active | Qwen-based / Apache 2.0 | Hugging Face, providers, self-host |
| Kimi K2.7 Code | Strongest commercially-licensed open coder | +21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active | 256K / Modified MIT | Hugging Face, DeepInfra, providers |
| DeepSeek V4-Pro | Agentic real-world work | II 44 (v4.1), 1.6T/49B active | 1M / MIT | DeepSeek API ($0.435/$0.87), local |
| Hy3 | Newest permissive-licence entrant | II 41 (v4.1); #22 on Arena WebDev (1,516.4) | Apache 2.0 | Hugging Face, providers, self-host |
| Inkling | Thinking Machines' first open model | II 41 (v4.1), agentic 32.3, released July 15 | Open weights | Hugging Face, providers, self-host |
| Inkling Small | Same family at a quarter the size | II 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio in | Apache 2.0 | Hugging Face (BF16 and NVFP4), providers, self-host |
| NVIDIA Nemotron 3 Ultra | NVIDIA-tuned, fully permissive license | II 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active | 1M / OpenMDW | OpenRouter, Hugging Face, AWS (8x B200 self-host) |
| Qwen 3.5 (397B / 17B active) | Multimodal, fast decode | 88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v6 | 1M / open | Together, OpenRouter, local |
| Qwen3.6-35B-A3B | Efficient open agentic coder (3B active) | 86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active | 262K (to 1M YaRN) / Apache 2.0 | Hugging Face, OpenRouter, local |
| Qwen3.6-27B | Laptop-runnable dense coder | 87.8 GPQA Diamond, dense 27B, multimodal | 256K / Apache 2.0 | Local Mac/PC, Hugging Face, OpenRouter |
| Rio 3.5 Open 397B | Qwen 3.5 fine-tune, multilingual reasoning | 70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5 | 397B/17B active / MIT | Hugging Face, providers, self-host |
| Llama 4 Maverick | Meta-line flagship | 17B active / 400B total params | Llama 4 license | Meta cloud, Hugging Face, local |
| NVIDIA Nemotron 3 Nano Omni | Edge / low-power | Multimodal, very small footprint | Compact / open | Local, NVIDIA tool |
Licensing matters as much as raw score here, and August widened the gap between the two groups. Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM 5.3 Flash (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Hy4 Preview (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use with no revenue test. The rest carry conditions worth reading before you build on them: MiniMax M3 and MiniMax H3 ship under their own community licences, H3 additionally requiring an application from the US, EU, UK and South Korea; Kimi K3 sits on a custom licence with a revenue threshold for anyone reselling it as a service; GLM 5.3 requires a Z.ai security review of any model-as-a-service operator above $10 billion in revenue; and the Qwen Community License 1.0 on Qwen3.8-Flash-Next requires a separate licence from Qwen to run a model-as-a-service or an AI work-assistant business, though internal use is exempt. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.
Benchmarks, Prices, and Hands-On Use
Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 2.1 board, covering the Intelligence and Agentic indexes, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.
Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We do not move a category crown on the strength of a single leaderboard, and where a new model is too recent for the human-preference boards to have rated it, we say so rather than crowning it on benchmark scores alone. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.
Fello AI brings the top AI models together in one lightweight native app.
Switch between them anytime, no separate subscriptions.
Download Fello AI




