Best AI Models in 2026
Rankings, comparisons, and deep dives, updated monthly as new models ship.
OpenAI launched GPT-6 Astra on September 3, the biggest launch of the month and not the biggest score. Artificial Analysis measures it at 61 on its Intelligence Index, exactly level with GPT-5.6 Sol and five points behind Claude Fable 5.1, at $10 / $50 per 1M tokens against Sol's $4 / $20. Where it does move is agentic coding: 67 on Artificial Analysis's Coding Agent Index running in Codex, level with Claude Opus 5 and Claude Fable 5 and behind Fable 5.1's 70, reached on a third of Sol's tokens and a fifth of Opus 5's. It is the first model OpenAI has designated Critical for cybersecurity, so approved defenders in its Daybreak program went first; ChatGPT Business and Pro followed on September 4 and Plus a few hours later, leaving the API, AWS, Azure and the free tier still to come. This page ranks on measured scores, so a model takes a crown as soon as it out-scores the holder on the boards that define the category, without waiting for the vote-based boards to catch up. A model nobody can use yet is flagged on the card rather than denied one.
Claude Fable 5.1 arrived on September 1 and went straight to the top of Artificial Analysis's Intelligence Index at 66, the highest score the house has measured, taking its Agentic Index at 61 with it. Claude Opus 5 becomes the value pick on the thing most readers are actually choosing on, the price, $5 / $25 per 1M tokens against Fable 5.1's $10 / $50. That crown is narrower than it was: Alibaba's Qwen3.8-Max-0902 debuted at #1 on Arena's Code Arena: WebDev board on September 2 with 1691 points against Opus 5's 1688, so Opus 5 now leads image-to-WebDev and sits second on WebDev. No Arena board has votes for 5.1 yet, so nothing this page decides on human preference has moved. Grok 4.6 was August's real surprise: a post-training refresh of Grok 4.5 that jumped five points to 61, tying GPT-5.6 Sol while charging $2 / $6, and finishing long agentic jobs in roughly half the turns Opus 5 needs. Meta is climbing faster than anyone: Muse Spark 1.3 landed on September 2 at 61 on the same index, eight points up from July, though it is a developer release with no consumer surface that loses every agent row on Meta's own chart. Our guide to AI benchmarks explains how that index is built.
The other move is at the cheap end, and for the first time this year it went the wrong way. DeepSeek raised both V4 models roughly fourfold on August 16 and split the rate card into peak and off-peak tiers, so DeepSeek V4-Flash 0731 now costs $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak instead of the flat $0.14 / $0.28 it launched on, which costs it the price-performance pick. That pick moves to GLM 5.3 Flash, which landed on August 26 with MIT weights on day one, an Intelligence Index of 57 and a list price of $0.15 / $0.50, halved to $0.075 / $0.25 until September 9, and which now undercuts DeepSeek at every hour of the day. Google had already halved its own Flash rate: Gemini 3.7 Flash shipped on August 13 at an introductory $0.75 / $3.75, with Gemini 3.6 Flash moved onto the same rate, and Gemini 3.8 Flash followed on September 2 at exactly that price, scoring 59 on Artificial Analysis's Intelligence Index against 56 for 3.7 Flash. All three prices double on January 1, 2027.
The open field then closed the month with three releases in three days: GLM 5.3 finally published its weights on August 28, though under a bespoke Z.ai licence rather than the plain MIT its smaller sibling got, Alibaba opened Qwen3.8-Flash-Next on August 26 as the first public look at the Qwen4 architecture, and Tencent open-sourced Hy4 Preview on August 28, 770 billion parameters under Apache 2.0. Below are the category winners for September 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks, and pricing are updated within 48 hours of any major model launch.
Want the top AI models without juggling separate subscriptions? Fello AI brings the leading models together in one native app for Mac, iPhone and iPad.
Download Fello AISep 3 OpenAI GPT-6 Astra ships at $10 / $50, ties GPT-5.6 Sol on the Intelligence Index at 61, and almost nobody can use it New
Sep 2 Alibaba Qwen3.8-Max-0902 takes Arena's WebDev board off Claude Opus 5 by three points, at $2 / $6 New
Sep 2 Muse Spark 1.3, Intelligence Index 57 to 61 at an unchanged $1.25 / $4.25, winning coding and losing every agent row New
Sep 2 Google Gemini 3.8 Flash, three points on the Intelligence Index at the same $0.75 / $3.75 New
Sep 1 Anthropic Claude Fable 5.1 and Mythos 5.1, a new #1 on Artificial Analysis at 66 and a 75% cut to cache reads New
Aug 28 Z.ai GLM 5.3 weights are out after the safety hold, but the licence is not MIT New
Aug 28 Tencent Tencent Hy4 Preview, 770B open weights under Apache 2.0 and the largest permissive model yet New
Aug 26 Z.ai GLM 5.3 Flash, Ox Alpha unmasked, and the first MIT weights Z.ai has shipped since GLM-5.2 New
Aug 26 Alibaba Qwen3.8-Flash-Next, open weights and the first public look at the Qwen4 architecture New
Aug 21 OpenAI GPT-5.6 Sol cut by more than 20%, now under Claude Opus 5 on both sides New
Aug 16 DeepSeek DeepSeek raises API prices roughly fourfold and splits the rate card into peak and off-peak tiers New
Aug 14 Z.ai GLM 5.3, a large agentic jump from post-training alone, shipped without the open weights New
Aug 13 Google Gemini 3.7 Flash, coding scores jump and the price halves to $0.75 / $3.75, until January New
Aug 12 SpaceXAI Grok 4.6, post-training refresh takes Intelligence Index 56 to 61 at an unchanged $2 / $6 New
Aug 5 Muse Spark 1.2 and Muse Code, Intelligence Index 51 to 57 at an unchanged $1.25 / $4.25, plus a terminal coding agent New
Aug 3 Alibaba Qwen 3.8 (Qwen3.8-Max), 2.4-trillion-parameter multimodal flagship now generally available at $2 / $6 New
Pending Google Gemini 3.5 Pro, still unreleased, months behind schedule per Bloomberg Delayed
Best AI for Writing
The best AI for writing is Claude Fable 5.1, which beats Claude Fable 5 on every board that has measured both, at the same $10 / $50, with Claude Sonnet 5 as the free-tier value pick and Claude Fable 5 as the pick if you weight Arena's human votes above the harness scores. Fable 5 leads Arena's creative-writing leaderboard and tops LiveBench Language at 90.7, but it has slipped to eighth on EQ-Bench Creative Writing v3, a board now led by GPT-6 Astra at 2163.9 with Claude Fable 5.1 second at 2152.7. This is a change from last month, when Claude Sonnet 5 held this slot on GDPval-AA; the preference boards do not support that placing, so Sonnet 5 is now the value pick rather than the quality leader. Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. If your writing is a work deliverable rather than prose, the GDPval-AA v2 professional-deliverables board is now led by Claude Fable 5.1 at 1853, ahead of Claude Opus 5 at 1824 and Fable 5 at 1723. Fable 5.1 landed on September 1 as the higher-capability model at the same $10 / $50, and it beats Fable 5 on every board that has measured both, including Humanity's Last Exam at 59.1% against 55.5%. We have moved the crown to it, because this page ranks on measured scores; Arena has not rated it yet, which is why Fable 5 stays in the table as the human-preference pick. Translation is a separate question with a separate winner, which we work through in our guide to the best AI for translation.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Fable 5.1 | Best writing overall | #1 Humanity's Last Exam (59.1%), #1 GDPval-AA v2 (1853), #1 Intelligence Index (66) | No Arena votes yet, so unrated on the preference boards | $10 / $50 |
| Claude Fable 5 | Top of Arena's human-vote boards | #1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #8 EQ-Bench | Priciest option here | $10 / $50 |
| Kimi K3 | Creative fiction and voice | #4 EQ-Bench Creative Writing (2070.6) | Only #10 on Arena creative writing | $3 / $15 |
| Claude Sonnet 5 | Free everyday writing | Free and default on claude.ai, 1M context | #53 Arena creative writing; 75.0 LiveBench Language | $2 / $10, now permanent |
| Claude Opus 5 | Professional deliverables on a budget | #2 GDPval-AA v2 at 1824, ahead of Fable 5 (1723) | Behind Fable 5 on Arena text, behind Fable 5.1 on GDPval-AA v2 | $5 / $25 |
| GPT-5.5 | Fact-anchored business writing | Documented factual-reliability gains over GPT-5.4 | Reasoning tiers now marked deprecated | $5 / $30 |
| Gemini 3.8 Flash | Bulk drafts at scale | Intelligence Index 59, 304.6 tok/s output, 1M context | Agent-tuned rather than prose-tuned; intro price doubles January 1, 2027 | $0.75 / $3.75 |
The writing crown moved to Claude Fable 5.1, and the rule behind this page moved with it: we now rank on measured scores rather than waiting for Arena's human votes. Fable 5.1 beats Fable 5 on every board that has measured both, taking Humanity's Last Exam at 59.1% against 55.5%, GDPval-AA v2 at 1853 against 1723 and the Intelligence Index at 66 against 62, all at the same $10 / $50. Claude Fable 5 keeps Arena's text and creative-writing boards, where it is still #1 at 1508.6 on the August 1 cutoff, so it stays in the table for anyone who weights human preference above harness scores, and Claude Sonnet 5 remains the free value pick. Two numbers we quoted last month were also re-fitted: AA-Omniscience now reads 43 for both Fable models rather than 40, and Fable 5 on Humanity's Last Exam reads 55.5% rather than 53.3%.
Best AI for Chat & Daily Assistant
The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena's text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $4 / $20 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| GPT-5.6 | Everyday chat, ChatGPT's default | The assistant most people can open; Terra matches GPT-5.5 at ~half cost | #14 on Arena text; scheming flagged by METR | Free / $20/mo Plus; API $0.20 / $1.20 to $4 / $20 |
| Claude Fable 5 | Highest-rated conversation | #1 on Arena text overall (1508.6) and 6 of 7 subcategories | No free tier; usage-credit access on Pro | $10 / $50 API |
| Claude Opus 5 | Thoughtful, nuanced answers | #1 Artificial Analysis Intelligence Index (63) and Agentic Index (55.3) | #6 on Arena text, behind Claude Fable 5 | $20/mo Pro, $5 / $25 API |
| GPT-5.5 Instant | Hallucination-sensitive daily work | 52.5% fewer hallucinated claims vs 5.3 Instant | Reasoning tiers now marked deprecated | $20/mo Plus; API $5 / $30 |
| Gemini 3.6 Flash | Fast, free, multimodal | Still the model the free Gemini tier serves, 1M context, #12 on Arena text | Weaker on hardest reasoning; 3.8 Flash outscores it at the same price | Free / $0.75 / $3.75 API |
| Fello AI | All the top models, one app | ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad | Routed via app, not direct | $9.99/mo |
GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open. OpenAI's GPT-6 Astra launched on September 3 and reached ChatGPT Business and Pro on September 4 and Plus a few hours later, but the pick stays with GPT-5.6 on reach: Astra went first to approved defenders in the Daybreak program, and OpenAI has announced nothing for the free tier.
Best AI for Images
The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena's text-to-image board at 1385 Elo and its image-editing board at 1463, and Artificial Analysis puts it first on its own image arena at 1339.4. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro. The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta's Muse Image is #3 on Arena's text-to-image board and #2 on image editing, the strongest showing any Meta image model has managed. Google's Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| ChatGPT Images 2.0 | Images with readable text | #1 text-to-image (1385) and #1 image editing (1463) | Less photoreal than the Gemini image line | Included in ChatGPT Plus |
| Reve 2.1 | Layout, typography, native 4K | #2 text-to-image at 1302, layout-preserving editing | Smaller ecosystem | Free / from $7.99/mo |
| Muse Image | Image editing, Meta ecosystem | #3 text-to-image, #2 image editing (1407) | New, thin tooling around it | Meta AI app |
| Nano Banana 2 (Gemini 3.1 Flash Image) | Photoreal portraits and products | Outranks Nano Banana Pro on both boards | Weaker on text in image | Gemini app / AI Studio |
| Seedream 5.0 Pro | Multilingual text + region editing | 10+ languages incl. Arabic RTL, lasso and layer editing | No independent benchmarks; copyright cloud | BytePlus / Magnific |
| Midjourney v8 | Stylized art, illustration | Aesthetic baseline most artists prefer | Weaker on text in image | $10-$120/mo |
| Grok Imagine | NSFW / Spicy Mode | Most permissive guardrails | Smaller model behind it | $30/mo SuperGrok |
The image crown is unchanged and GPT Image 2 still leads all three boards we track, on Arena text-to-image (1385), Arena image editing (1463) and Artificial Analysis's image arena (1339.4). The one addition is Microsoft's MAI-Image-2.5, which sits third on Artificial Analysis's image board at 1269.7 and third on Arena's image-editing board, so third place now depends on which house you read.
Best AI for Video
The best AI for video generation is Gemini Omni 1.1 Flash, which Google shipped on August 27 and which now leads Arena's text-to-video board at 1515, just ahead of its own predecessor Gemini Omni Flash at 1511. It is the more capable model on specification too: scene extension in 10-second increments to a cumulative 40 seconds, first and last frame control, video references, 360p drafts and 1080p or 4K output, priced at roughly $0.03 per second at 360p, $0.10 at 720p, $0.15 at 1080p and $0.30 at 4K. Gemini Omni Flash remains the cheaper option at about $0.10 a second and is a statistical tie on Arena, but it caps at 10-second clips. Neither leads Artificial Analysis's video board any more: Alibaba's Wan 3.0, an API-only model released on August 24 that builds video from documents, spreadsheets and slides, tops it at 1239 with Omni Flash on 1238 and MiniMax H3 Max on 1235. If you need longer takes inside Google's tooling, Veo 3.1 still runs in the Gemini app, AI Studio and Vertex AI with native audio and 1080p output. MiniMax H3 (July 31) remains the specification contender, with 2K clips of 4 to 15 seconds and native stereo audio from $0.13 a second.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Gemini Omni 1.1 Flash | Best AI video overall | #1 Arena text-to-video (1515); 40-second scene extension, 4K output | Second to Wan 3.0 on Artificial Analysis's board | $0.03-$0.30/sec; Gemini API / AI Studio / Flow |
| Gemini Omni Flash | Cheaper 10-second clips | #2 Arena text-to-video (1511), #2 AA video (1238); conversational editing | Caps at 10-second generations | ~$0.10/sec; Gemini app / AI Studio |
| Wan 3.0 | Video from documents and slides | #1 on Artificial Analysis's video board (1239); 2-30s, up to 1080p, Omni-Reference input | API-only, no published weights; #3 on Arena | $0.05-$0.20/sec |
| MiniMax H3 | 2K clips with native audio | 4-15s at 2K, native stereo audio; #8 Arena text-to-video (1462) | #4 on AA video (1227); the open weights exclude the 2K upscaler and need an application in the US, EU, UK and South Korea | $0.13/sec at 2K |
| Dreamina Seedance 2.5 | ByteDance challenger | #6 Arena text-to-video (1482); leads image-to-video with audio | ByteDance ecosystem, limited Western access | Dreamina / BytePlus |
| Muse Video | Meta ecosystem video | #3 text-to-video at 1459 | Newest of the group, thin tooling | Meta AI app |
| Veo 3.1 | Longer production clips | Native audio, 1080p, strong physics consistency | #6 on Arena video, not the quality leader | Google AI Pro / Ultra |
| Kling 3.0 / 3.0 Turbo | Fast iteration at lower cost | Native 4K, 60fps, 15-second clips; Turbo shipped June 17 | Outside the top 16 on Arena text-to-video | From $10/mo |
| Luma Ray 3 | Photoreal scenes | Strong realism for landscapes | Smaller community | Free / from $9.99/mo |
The video crown moved to Gemini Omni 1.1 Flash, which Google shipped on August 27 and which now leads Arena's text-to-video board at 1515 against Gemini Omni Flash's 1511. The two sit inside each other's confidence intervals, so read that as a tie broken on capability rather than a decisive win: 1.1 Flash extends scenes to a cumulative 40 seconds and outputs 4K, where Omni Flash caps at 10 seconds. The bigger change is that neither model leads Artificial Analysis's video board any more. Alibaba's Wan 3.0 took it on August 24 at 1239, ahead of Omni Flash at 1238 and MiniMax H3 Max at 1235, and it accepts documents, spreadsheets and slides as input. Wan 3.0 is API-only with no published weights, so it does not change our open-weight picks. MiniMax H3 also picked up Arena votes this month, entering the text-to-video board at #8 with 1462.
Best AI for Coding
The best AI for coding is Claude Fable 5.1, which leads the harness benchmarks at 55.8% on Terminal-Bench 4.0 and 70 on Artificial Analysis's Coding Agent Index, with Claude Opus 5 the value pick at half the price and still the leader on the boards where developers vote on the finished result. Alibaba's Qwen3.8-Max-0902 debuted at #1 on Arena's WebDev board on September 2 with 1691 points against Opus 5's 1688, a three-point margin, and Opus 5 keeps image-to-WebDev at 1,668.6. Price is the second half of the argument: Opus 5 runs $5 / $25 per 1M tokens against Claude Fable 5's $10 / $50, and Artificial Analysis measures it at $2.34 per index task against Fable 5's $3.14. Anthropic's own docs tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for workloads that need the highest available capability. Claude Fable 5 is the runner-up and stays the pick for the hardest long-horizon work, at 1,630.7 on WebDev and #2 on image-to-WebDev (1,625.7); its September 1 successor, Claude Fable 5.1, is the stronger model on the harness benchmarks at the same price, reaching 55.8% on Terminal-Bench 4.0 against Opus 5's 52.3% and Fable 5's 42.0% on Anthropic's own figures, but Arena has no votes for it yet. The contender is Kimi K3, now #3 on WebDev at 17 points off the lead and #3 on Arena's Agent board, the strongest open-weight coder on the boards, though self-hosting it means 1.56 TB of weights. On Artificial Analysis's Coding Index it is not a Claude sweep: GPT-5.6 Sol (xhigh) leads at 78.3 with Opus 5 (max) at 78.0, close enough to call a tie. The cheapest serious contender is now GLM 5.3 Flash at $0.15 / $0.50 under MIT, after DeepSeek's August 16 rate rise took V4-Flash 0731 to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak, and the best value at the frontier is Grok 4.6, which posts 88.4% on Terminal-Bench v2.1 and $0.84 per index task at $2 / $6.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Fable 5.1 | Best coding overall | #1 Intelligence Index (66) and #1 Agentic Index (61); 55.8% Terminal-Bench 4.0 (Anthropic) | No Arena votes yet; twice the price of Opus 5 | $10 / $50 |
| Claude Opus 5 | Best value at half the price | #1 Arena image-to-WebDev (1,668.6) and #2 on WebDev (1688); $2.34 per index task | Lost Arena's WebDev board to Qwen3.8-Max-0902 on September 2; Artificial Analysis Coding Index has GPT-5.6 Sol a shade ahead | $5 / $25 |
| Qwen3.8-Max-0902 | Top of Arena's WebDev board | #1 Code Arena: WebDev (1691), three points above Opus 5; 2.4T parameters, 1M context | QwenCloud API only with no consumer front-end; no independent composite score yet, and three points is inside the noise | $2 / $6 |
| GPT-6 Astra | Token-efficient agentic coding | 67 on Artificial Analysis's Coding Agent Index in Codex, level with Opus 5 and Fable 5; one third of GPT-5.6 Sol's tokens and one fifth of Opus 5's | Not on ChatGPT's free tier; twice Opus 5's price | $10 / $50 |
| Claude Fable 5 | Hardest long-horizon agentic work | #2 Arena image-to-WebDev (1,625.7), #4 WebDev; 1M context | Priciest; Artificial Analysis Coding Index puts it 7th | $10 / $50 |
| Kimi K3 | Web app building and agents | #3 Arena WebDev, #3 Arena Agent, highest open Intelligence Index (60) | 1.56 TB to self-host | $3 / $15 |
| GPT-5.6 Sol | OpenAI flagship, agentic coding | #1 Artificial Analysis Coding Index at 78.3 (xhigh) | Absent from the official Terminal-Bench board; eval-gaming flagged by METR | $4 / $20 |
| Muse Spark 1.3 | Cheap agentic coding | Intelligence Index 61 at $1.25 / $4.25; wins all three coding rows on Meta's own chart, 88.8% on Terminal-Bench 2.1 | Meta's own numbers, measured on the max variant that is a partners-only preview; loses all six agent rows on the same chart, and Scale SEAL has rated only 1.1 | $1.25 / $4.25 |
| Grok 4.6 | Cheap value coder | 88.4% on Terminal-Bench v2.1 and Intelligence Index 61, at $0.84 per index task | Prompts of 200K tokens and up re-bill the whole request at double; higher hallucination rate | $2 / $6 |
| Gemini 3.8 Flash | Agent coding at scale | Intelligence Index 59, 73.7% DeepSWE v1.1, 89.4% Terminal-bench 2.1, 304.6 tok/s | 19.1% on Terminal-bench 4.0 against Opus 5's 51.8%; intro price doubles January 1, 2027 | $0.75 / $3.75 |
| GLM 5.3 Flash | Best open-weight coder you can host | Intelligence Index 57; DeepSWE v1.1 63.4 and Terminal Bench 2.1 84.3 (vendor figures); MIT, 320B/18B active | Kimi K3 outscores it; the coding figures are vendor-reported and it has no Arena votes yet | Open weights (MIT) |
The coding crown moved to Claude Fable 5.1 under the same rule change. It leads the harness benchmarks at 55.8% on Terminal-Bench 4.0 against Claude Opus 5's 52.3% and Claude Fable 5's 42.0%, and takes Artificial Analysis's Coding Agent Index at 70 against 67 for both Opus 5 and GPT-6 Astra. Claude Opus 5 becomes the value pick and keeps a real argument: $5 / $25 against Fable 5.1's $10 / $50, $2.34 per index task, and Anthropic's own docs still tell developers to start with it. The vote-based boards are now split: Alibaba's Qwen3.8-Max-0902 took Arena's WebDev board on September 2 with 1691 against Opus 5's 1688, Opus 5 keeps image-to-WebDev at 1,668.6, and Arena has no votes for Fable 5.1 at all. GPT-6 Astra (September 3) reaches 67 on the Coding Agent Index on one fifth of Opus 5's tokens, the largest efficiency gain published this month, but it lists at twice Opus 5's price and is not on ChatGPT's free tier. Kimi K3 stays the strongest open-weight coder, with GLM 5.3 Flash the one to self-host at Intelligence Index 57 under MIT.
Best AI for Creativity
The best AI for unfiltered, on-trend creative work is Grok 4.6, and we want to be exact about why. This pick is about the product, not the prose quality. Grok carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, and the boards were blunt about its predecessor: Grok 4.5 sat #33 on EQ-Bench Creative Writing and #41 on Arena's creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. Neither creative-writing board has rated Grok 4.6 yet, and since SpaceXAI aimed this release at coding and agentic work rather than prose, we are not assuming it moved. If you are picking on output quality alone, Claude Fable 5 still leads Arena's creative-writing board, while EQ-Bench now puts GPT-6 Astra first at 2163.9, Claude Fable 5.1 second and Kimi K3 fourth. Choose Grok 4.6 for what it will let you make, not for how well it writes.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Grok 4.6 | Unfiltered, opinionated, on-trend | Fewest content restrictions, native real-time X grounding | Creative-writing boards have not rated it; Grok 4.5 sat #33 EQ-Bench, #41 Arena | $30/mo SuperGrok |
| Claude Fable 5 | Highest-quality creative prose | #1 Arena creative writing, #1 LiveBench Language | Cautious guardrails on edgy briefs | $10 / $50 API |
| Kimi K3 | Fiction and distinctive voice | #4 EQ-Bench Creative Writing at 2070.6 | Only #10 on Arena creative writing | $3 / $15 |
| Claude Opus 5 | Long-form structured creativity | Holds long threads and self-edits; Intelligence Index 63, second only to Claude Fable 5.1 | Most cautious of the group | $20/mo Pro; $5 / $25 API |
| Gemini 3.1 Pro | Multimodal creative | Strong text, image and video chain | Quotas inside the Gemini app | Free / $2.00-$4.00 API in |
| Grok Imagine (Spicy Mode) | NSFW / adult creative | Most permissive image generation | Niche use case | $30/mo SuperGrok |
Grok keeps this pick, now on Grok 4.6, which replaced Grok 4.5 as SpaceXAI's flagship on August 12. Nothing about the product's permissiveness or its live X access changed, which is what the pick rests on. The model underneath is meaningfully stronger, up five points to Intelligence Index 61, but that gain landed in coding and agentic work rather than prose, and no creative-writing board has rated 4.6 yet. Claude Fable 5 remains the model to use when you want the better writing.
Best AI for Accuracy & Research
The best AI for accuracy and research is Claude Fable 5.1, which holds the highest factual accuracy Artificial Analysis has measured, 67% on AA-Omniscience, and tops Humanity's Last Exam at 59.1%. The value pick is Gemini 3.1 Pro, and it is still what we would reach for when cost matters. Its strongest result is on ARC Prize's ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you want when the answer has to be current rather than merely plausible. It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity's Last Exam, and scores 46.44 on Scale SEAL's HLE board, which Claude Fable 5 now leads at 55.5%. We have dropped the previous ARC-AGI-2 framing: its 77.1% is still correct, but the board has moved and that score now places it around 14th. Two honest caveats. On grounded search specifically, Arena's search leaderboard is led by OpenAI rather than by Google or Anthropic: GPT-5.6 Sol tops it at 1257, with Claude Fable 5 fifth and Gemini 3.1 Pro grounding ninth. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience, a board Claude Fable 5.1 now shares the top of.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| Claude Fable 5.1 | Highest measured factual accuracy | 67% factual accuracy on AA-Omniscience, the highest Artificial Analysis has measured; #1 Humanity's Last Exam (59.1%) | No cheap tier, and unrated on Arena's search leaderboard | $10 / $50 |
| Gemini 3.1 Pro | Best value, cheap factual work | 98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol) | ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search | $2.00-$4.00 / $12.00-$18.00 (tiered) |
| Claude Fable 5 | Grounded search | #5 on Arena's search leaderboard, behind GPT-5.6 Sol | No single cheap tier | $10 / $50 (Fable 5) |
| GPT-5.6 Sol | Novel reasoning | #1 ARC-AGI-2 at 92.5%, against a 100% human panel | Scheming flagged by METR | $4 / $20 |
| Claude Opus 5 | Hardest unseen problems | #1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize) | Artificial Analysis measures a 50% hallucination rate | $5 / $25 |
| Qwen 3.7 Max | Frontier accuracy at value pricing | 92.4 GPQA Diamond, 200 free requests/day | API-only, no chat front-end | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Opus 4.6 | Honesty under pressure | #1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5 | Superseded as a flagship | Legacy Anthropic model |
The accuracy crown moved to Claude Fable 5.1, which holds the highest factual accuracy Artificial Analysis has measured, 67% on AA-Omniscience against Claude Fable 5's 65%, and tops Humanity's Last Exam at 59.1% against 55.5%. The AA-Omniscience index itself reads 43 for both models, up from the 40 we quoted last month, because 5.1 buys its extra accuracy with a higher attempt rate on questions it cannot answer. Gemini 3.1 Pro becomes the value pick and is still what we would reach for on cost: 98% on ARC-AGI-1 at $0.52 a task, tying the human panel, plus native Google Search grounding, though its GPQA Diamond re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly. GPT-6 Astra (September 3) is the launch to watch here rather than a new crown: Artificial Analysis measures its hallucination rate falling from Sol's 92% to 51% at max effort while accuracy rises four points, the largest honesty improvement any vendor has posted this year, but it loses about 80 Elo on GDPval-AA v2 and is not on ChatGPT's free tier.
Best AI for Problem Solving
The best AI for hard problem solving is GPT-6 Astra, which took the reasoning boards off GPT-5.6 Sol within days of launching. It leads LiveBench Reasoning at 92.65 and ARC-AGI-2 at 95%, the highest anyone has scored against that board's 100% human panel, and sits second on LiveBench Mathematics at 96.81 behind Claude Fable 5.1's 97.01. It also reports 97.6% on FrontierMath Tier 4 v2, a number worth carrying with Epoch AI's disclosure that OpenAI funded the benchmark and holds exclusive access to part of it. The catch is price and reach: $10 / $50 per 1M tokens against Sol's $4 / $20, and it needs a paid ChatGPT plan. GPT-5.6 Sol is the value alternative, third on both LiveBench boards at 96.20 and 91.65 and second on ARC-AGI-2. Qwen 3.7 Max is the budget pick for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, with 200 free model requests a day. Claude Opus 5 remains the alternative for long agentic reasoning chains, though Astra has taken ARC-AGI-3 from it as well.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| GPT-6 Astra | Hardest math, science and reasoning | #1 LiveBench Reasoning (92.65), #1 ARC-AGI-2 (95%), #1 ARC-AGI-3 | Needs a paid ChatGPT plan; 2.5x Sol's price | $10 / $50 |
| Claude Opus 5 | Long agentic reasoning chains | #2 AA-Briefcase (1645) and #2 GDPval-AA v2 (1737); 30.2% ARC-AGI-3 | Astra now leads ARC-AGI-3; 50% hallucination rate | $5 / $25 |
| GPT-5.6 Sol | Value STEM flagship | #2 ARC-AGI-2 (92.5%); #3 LiveBench Mathematics (96.20) and Reasoning (91.65) | FrontierMath still unpublished; scheming flagged by METR | $100/mo ChatGPT Pro; API $4 / $20 |
| Qwen 3.7 Max | Competition math on a budget | 97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day | API-only | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Fable 5 | Math inside a coding workflow | #1 on Arena's math subcategory (1543), 96.0 LiveBench Mathematics | Priciest option here | $10 / $50 |
| GLM 5.3 Flash | Open-weight problem solving | Intelligence Index 57 under MIT, four points above GLM-5.2 at less than half the size; 320B/18B active, 1M context | Needs a multi-GPU server, not a workstation; GLM 5.3 scores higher on Z.ai's own figures and is downloadable since August 28, but at more than twice the size and under a bespoke licence | Open weights (MIT) |
The problem-solving crown moved to GPT-6 Astra, which took the boards this category is decided on within days of its September 3 launch. It leads LiveBench Reasoning at 92.65 against GPT-5.6 Sol's 91.65, takes ARC-AGI-2 at 95% against Sol's 92.5%, and displaced Claude Opus 5 on ARC-AGI-3, where its 62.7% on the standard harness beats Opus 5's 30.2%. Read the ARC-AGI-3 figures carefully: that board measures human-level action efficiency on unseen environments rather than problems solved, and the widely quoted 98.6% comes from a provider-adapter harness rather than the standard one. Sol drops to third on both LiveBench boards but stays the value alternative at $4 / $20 against Astra's $10 / $50, and Claude Fable 5.1 quietly took LiveBench Mathematics at 97.01.
Best AI Agent
The best AI agent is Gemini Spark if you want the agent in the cloud and Claude Cowork if you want it on your desktop. These are joint picks rather than a first and a second: Spark runs in a Google Cloud VM, so it keeps working while your laptop is shut, wired into Gmail, Docs and, since early September, Google Photos; Cowork runs on your Mac or Windows machine and drives your local apps and your screen. No independent board rates the two products against each other, so the choice is about where the work has to happen rather than which one scores higher. Price is no longer a tiebreaker either: Spark now reaches the $19.99/month Google AI Pro tier in the US and more than 160 other countries, and Cowork is included at no extra charge on every paid Claude plan from $20/month. ChatGPT Codex Mobile (May 14) is the alternative for coding-agent work, and OpenAI Operator-class browser agents for web tasks. Read the full Gemini Spark vs Claude Cowork comparison.
| Agent | Best For | Where It Runs | Strength | Price |
|---|---|---|---|---|
| Gemini Spark | 24/7 cloud tasks, Workspace workflows | Google Cloud VM (always-on) | First true 24/7 agent, deep Workspace integration | From $19.99/mo Google AI Pro |
| Claude Cowork | Desktop, app-driving, design + code | Your Mac/Windows desktop | Drives local apps, sees your screen | $20/mo Claude Pro |
| ChatGPT Codex Mobile | Coding agent on phone | OpenAI cloud + iOS/Android | Approve diffs and redirect work from phone | Included in ChatGPT plans |
| Grok Agentic (Grok 4.6) | Real-time research, X scraping | SpaceXAI cloud | Native X integration; GDPval-AA v2 Elo 1665 | $30/mo SuperGrok |
| OpenAI Operator-class | Browser tasks, web forms | OpenAI cloud + your browser | Web automation | ChatGPT Pro |
No new consumer agent product shipped, though Gemini Spark did expand: it reached the $19.99/month Google AI Pro tier and gained Google Photos, where it can search, edit, curate and run scheduled workflows. Meta's Muse Code (August 5) is a terminal developer tool rather than a consumer one, so the Spark (cloud) versus Cowork (desktop) choice still drives most agent decisions. The model layer underneath moved a long way. Claude Fable 5.1 (September 1) now leads Arena's Agent board outright, taking it off Claude Opus 5, which sits second and third on its two effort settings, and it also tops Artificial Analysis's AA-Briefcase at 1662 against Opus 5's 1645 and GDPval-AA v2 at 1766 against 1737. Two housekeeping notes on those numbers: Artificial Analysis retired the standalone Agentic Index with Intelligence Index v4.2 on September 4 and now measures agentic work through AA-Briefcase and GDPval-AA v2 instead, and it re-anchored the GDPval-AA v2 Elo scale at the same time, which is why these figures sit roughly 90 points below the ones we quoted last month. Opus 5 is still the one most teams should build on at $5 / $25, half Fable 5.1's price. Meta's Muse Spark 1.3 (September 2) deserves a fairer hearing than its own launch chart gives it: Meta groups six benchmarks under Agent and 1.3 loses all six, but on Artificial Analysis's independent GDPval-AA v2 board it places third among distinct models at 1719, ahead of GLM 5.3, Grok 4.6 and Claude Fable 5, at $1.25 / $4.25. It also wins long context, 98.1 on MRCR's 512K to 1M band against 55.5 for 1.2. Scale SEAL has still rated only the older 1.1, which leads its MCP Atlas tool-use board at 88.1. GLM 5.3 Flash (August 26, MIT) is the open-weight agent model we would host, at AutomationBench 48.8 on Z.ai's own chart where Claude Opus 4.8 scores 41.0; on Arena's Agent board GLM-5.2 now sits twelfth and Kimi K3 seventh. Grok 4.6 remains the cost story rather than a new agent product: a GDPval-AA v2 Elo of 1665, sixth among distinct models, while finishing long-horizon jobs in roughly 53 turns and 0.5 billion input tokens against about 103 turns and 2.0 billion for Opus 5. GPT-6 Astra (September 3) argues for Codex rather than for a new agent product, at 67 on the Coding Agent Index against Fable 5.1's 70 and on a fifth of Opus 5's tokens.
The leading models like ChatGPT, Claude and Gemini, together on Mac, iPhone and iPad.
Free to start, 4.7★ across 27,000+ reviews.
Download Fello AIBest AI for Students
The best AI for students is GPT-5.6 Luna inside ChatGPT for general coursework and Gemini 3.6 Flash inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Opus 5 as the alternative for essay editing. Most students don't need to pay, and the free tier improved on August 6: GPT-5.6 Luna became the default for ChatGPT Free and Go, with unlimited text chats and a Think button for harder questions, though file uploads and image tools stay capped. Gemini 3.6 Flash is still what the free Gemini app serves, Claude Sonnet 5 is the free Claude default, and DeepSeek V4 is free on DeepSeek's chat site. Luna is the cheapest member of the GPT-5.6 family rather than the strongest, so reach for the Think button or switch to GPT-5.5 when an essay needs more care. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's paid flagship and GPT-6 Astra now reports 97.6% on FrontierMath Tier 4 v2 against GPT-5.5 Pro's verified 39.6%, though Astra is not on ChatGPT's free tier; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).
| Task | Best Model | Why | Free? | Alternative |
|---|---|---|---|---|
| Essays & coursework | GPT-5.6 Luna | Free default in ChatGPT since August 6; unlimited text chats | Yes | Claude Sonnet 5 (free Claude) |
| STEM problem-solving | GPT-5.6 Sol / Qwen 3.7 Max | Paid STEM flagship; Astra reports 97.6% FrontierMath Tier 4 v2 / 97.1 HMMT 2026 Feb | Pro paid / Qwen API paid | Gemini 3.6 Flash (free) |
| Research & accuracy | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, native Google Search grounding | Yes (Gemini app) | Claude Opus 5 |
| Writing editing | Claude Sonnet 5 | Free and default on claude.ai; Claude Fable 5 is the quality leader | Yes (Claude free) | Claude Fable 5 |
| Multimodal study (PDFs, slides, images) | Gemini 3.6 Flash | 1M context, free in Gemini app | Yes | NotebookLM (Google) |
The free tier is where this guide moved. On August 6 OpenAI made GPT-5.6 Luna the default for ChatGPT Free and Go and gave both unlimited text chats plus a Think button for harder questions, so the free pick is now Luna rather than GPT-5.5, which stays selectable when an essay needs more care. File uploads, images and other tools are still capped. Nothing moved on the Google side: Gemini 3.8 Flash shipped on September 2 but reaches free users only through AI Studio and the API, so the free Gemini app still serves 3.6 Flash and it keeps the multimodal row. GPT-6 Astra (September 3) is the one to watch for hard problem sets, reporting 97.6% on FrontierMath Tier 4 v2 against GPT-5.5 Pro's verified 39.6%, but it needs a paid ChatGPT plan and no free tier includes it.
Best AI for Work & Professionals
The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark is the only true 24/7 cloud agent, and since July 30 it reaches the $19.99/month Google AI Pro tier as well as Google AI Ultra.
| Use Case | Best Model | Key Stat | Price | Alternative |
|---|---|---|---|---|
| Daily knowledge work | GPT-5.6 | ChatGPT's default since July 9; the assistant most people can open | $20/mo ChatGPT Plus | Claude Opus 5 |
| Coding (proprietary) | Claude Opus 5 | #1 on Arena image-to-WebDev (1,668.6) and #2 on WebDev (1688); Anthropic's recommended default | $20/mo Claude Pro | Claude Fable 5 |
| Coding (cost-effective) | Qwen3.8-Max-0902 | #1 Code Arena: WebDev (1691), three points above Claude Opus 5; API-only on QwenCloud | $2 / $6 | Qwen 3.7 Max; DeepSeek V4-Flash 0731 |
| Research & briefings | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, Google grounding | Google AI Pro / Ultra | Claude Opus 5 |
| Hard math, physics, finance modelling | GPT-6 Astra | #1 LiveBench Reasoning and ARC-AGI-2 (95%); reports 97.6% FrontierMath Tier 4 v2 | $100/mo ChatGPT Pro | GPT-5.6 Sol; Qwen 3.7 Max |
| Always-on agent workflows | Gemini Spark | First 24/7 cloud agent | From $19.99/mo Google AI Pro | Claude Cowork |
| Live news, X-context creative | Grok 4.6 | Intelligence Index 61 + native X grounding | $30/mo SuperGrok | Gemini 3.1 Pro |
| All-in-one consolidation | Fello AI | ChatGPT + Claude + Gemini + Grok + DeepSeek | $9.99/mo | Pay each vendor separately |
Three launches landed in the first three days of September and none of them moves the daily-work pick. Claude Fable 5.1 (September 1) tops Artificial Analysis's Intelligence Index at 66 and its Agentic Index at 61, but Claude Opus 5 stays the recommendation here at $5 / $25, half Fable 5.1's price, and Anthropic's own docs still tell teams to start with Opus 5. Alibaba's Qwen3.8-Max-0902 (September 2) took Arena's WebDev board from Opus 5 by three points and takes the cost-effective coding row at $2 / $6, though it is API-only on QwenCloud with no consumer front-end. GPT-6 Astra (September 3) is the launch that could move the daily-work row, and it has not yet: Astra reached Business and Pro seats on September 4 and Plus a few hours later, but at $10 / $50 against Sol's $4 / $20 it is a costlier default, so GPT-5.6 keeps the row.
AI Model Pricing in September 2026
From $0 free tiers to $199.99/month Google AI Ultra. The most consequential price on this table is Claude Opus 5 at $5 / $25, because it is the second-highest scorer on Artificial Analysis's Intelligence Index at half the cost of the two Fable models around it. Claude Fable 5.1, which took the top of that index on September 1, holds the same $10 / $50 list price as Fable 5 but cuts cache reads 75% to $0.25, which is where the saving on long agentic runs actually sits rather than on the sticker. Meta's Muse Spark 1.3 lists at $1.25 / $4.25, unchanged across three capability bumps since July, with a Contributor tier at $0.10 / $0.20 for anyone willing to let Meta train on their prompts, and Grok 4.6 lists at $2 / $6, with the caveat that a prompt of 200,000 tokens or more re-bills the entire request at $4 / $12. The price-performance pick moved in August, and not because anything got cheaper: DeepSeek raised both V4 models roughly fourfold on August 16, taking V4-Flash 0731 from a flat $0.14 / $0.28 to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak. The pick is now GLM 5.3 Flash at $0.15 / $0.50 list, halved until September 9, which scores 57 on the Intelligence Index and undercuts even DeepSeek's off-peak rate on both sides at every hour of the day. At the frontier itself the value pick is Grok 4.6, scoring 61 for $0.84 per index task where Opus 5 costs $2.34 and GPT-5.6 Sol charges $4 / $20 for the same score after its August 21 cut. Among closed models GPT-5.6 Luna is the cheapest at $0.20 / $1.20 after OpenAI's July 30 cut. For a deeper breakdown see our full AI Pricing Comparison Guide. The month's new flagship does not change that shape: GPT-6 Astra arrived on September 3 at $10 / $50, the same list price as Claude Fable 5.1 and 2.5 times GPT-5.6 Sol, with cached input at $1 and cache writes at $12.50, and Artificial Analysis measures it 75% more expensive per index task than Sol for the same score of 61.
| Model | Input (per 1M) | Output (per 1M) | Context | Free Access? |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 | 1,050,000 (922,000 max input) | ChatGPT Business/Pro since September 4, Plus hours later; no free tier. API, AWS and Azure still to come |
| GPT-5.5 | $5.00 | $30.00 | 1M (400K in Codex) | ChatGPT Free; API paid |
| GPT-5.5 Pro | $30.00 | $180.00 | 1M | ChatGPT Pro from $100/mo ($200 higher-usage tier) |
| GPT-5.6 Sol | $4.00 | $20.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Terra | $2.00 | $12.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Luna | $0.20 | $1.20 | Not published | Live in ChatGPT, Codex & API (July 9) |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Claude Pro/Max default; API paid |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Legacy model at Anthropic; Pro/Max/API |
| Claude Fable 5.1 | $10.00 | $50.00 | 1M | Cache reads $0.25; in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits. Mythos 5.1 bills identically, by invitation only |
| Claude Fable 5 | $10.00 | $50.00 | 1M | Permanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits |
| Claude Sonnet 5 | $2.00 | $10.00 | 1M | Claude Free & Pro default; API paid |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | API paid (superseded by Sonnet 5) |
| Gemini 3.1 Pro | $2.00 (≤200K) / $4.00 (>200K) | $12.00 (≤200K) / $18.00 (>200K) | 1M | Limited Gemini app; API paid |
| Gemini 3.8 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | AI Studio, Antigravity, Gemini Enterprise + paid API; in the Gemini app reported for AI Pro/Ultra |
| Gemini 3.7 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | AI Studio, Android Studio, Antigravity + paid API; in the Gemini app via Spark only (AI Pro/Ultra, excludes EEA/UK/CH/Nigeria) |
| Gemini 3.6 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | Free Gemini app default; AI Studio; free API tier + paid API |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | AI Studio; free API tier + paid API |
| Qwen3.8-Max-0902 | $2.00 ($0.17 explicit cache hit, $0.25 implicit) | $6.00 | 1M (128K output) | QwenCloud API only; no consumer chat front-end, no open weights |
| Qwen 3.7 Max | $1.25 promo / $2.50 list | $3.75 promo / $7.50 list | 1M | 200 free requests/day; API paid beyond that |
| MiniMax M3 | $0.30 (50% off $0.60) | $1.20 (≤512K) | 1M | Open weights; hosting costs apply |
| LongCat-2.0 | Provider-dependent | Provider-dependent | 1M | Open weights (MIT); hosting costs apply |
| NVIDIA Nemotron 3 Ultra | Provider-dependent | Provider-dependent | 1M | Open weights (OpenMDW); hosting costs apply |
| Qwen 3.5 (open-weight) | Self-host / Together | Self-host / Together | 1M | Open weights; hosting costs apply |
| Nex-N2-Pro | Self-host / providers | Self-host / providers | 1M | Open weights (Apache 2.0); hosting costs apply |
| Rio 3.5 Open 397B | Self-host / providers | Self-host / providers | 1M | Open weights (MIT); hosting costs apply |
| Grok 4.3 | $1.25 | $2.50 | 1M | Free consumer plan; API paid |
| Grok 4.6 | $2.00 (<200K) / $4.00 (≥200K) | $6.00 (<200K) / $12.00 (≥200K) | 500K | SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel, Cloudflare |
| Muse Spark 1.3 | $1.25 ($0.10 Contributor tier) | $4.25 ($0.20 Contributor tier) | 1M | Paid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount |
| Kimi K3 | $3.00 ($0.30 cache-hit) | $15.00 | 1M | Free basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License) |
| Gemini Omni Flash (video) | $1.50 | $17.50 (video output) | 10-second clips | Gemini app / Flow; AI Studio + API |
| DeepSeek V4-Pro | $0.66 off-peak / $1.32 peak ($0.022 cache-hit off-peak) | $1.98 off-peak / $3.96 peak | 1M | DeepSeek Chat free; API paid, peak/off-peak tiers since August 16 |
| DeepSeek V4-Flash 0731 | $0.22 off-peak / $0.44 peak ($0.007 cache-hit off-peak) | $0.66 off-peak / $1.32 peak | 1M | DeepSeek Chat free; API paid, peak/off-peak tiers since August 16 |
| Kimi K2.7 Code | Provider-dependent | Provider-dependent | 256K | Open weights; hosting costs apply |
| GLM-5.2 | Provider-dependent | Provider-dependent | 1M | Open weights; hosting costs apply |
| GLM 5.3 | $1.40 ($0.26 cache-hit) | $4.40 | 1M | Z.ai API, GLM Coding Plan and ZCode; weights on Hugging Face since August 28 under the bespoke GLM 5.3 License |
| GLM 5.3 Flash | $0.15 ($0.03 cache-hit); $0.075 promo | $0.50; $0.25 promo | 1M | Z.ai API, GLM Coding Plan, OpenRouter; MIT weights on Hugging Face; promo ends September 9 |
| Tencent Hy4 Preview | $0.834 ($0.042 cache-hit) | $2.501 | 1M+ | Open weights (Apache 2.0); Tencent Cloud TokenHub, OpenRouter, WorkBuddy, CodeBuddy |
| Qwen3.8-Flash-Next | Provider-dependent | Provider-dependent | 262K (to 1M) | Open weights (Qwen Community License 1.0); hosting costs apply |
| ERNIE 5.1 | China-region pricing | China-region pricing | 256K | Baidu free tier |
| Gemini Spark (agent) | Not API-priced | Not API-priced | 1M (Gemini base) | Google AI Pro $19.99, AI Ultra $99.99 or $199.99/mo |
| Fello AI (aggregator) | Routed via app | Routed via app | Model-dependent | $9.99/mo, free tier available |
The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. Grok 4.6 works the same way but bites harder: at 200,000 prompt tokens and above, SpaceXAI re-bills the entire request at the higher rate rather than only the overage, so crossing the line by a thousand tokens doubles the cost of the whole call. DeepSeek is the other rate card that needs reading twice: since August 16 both V4 models bill at peak rates from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays, and at exactly half that in every other hour, so the same job can cost twice as much depending on when you run it. If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.
Best Open-Weight Models in September 2026
The best open-weight model in September 2026 is Kimi K3, and it took the lead the moment Moonshot published the weights on July 27, 2026. It holds the highest Intelligence Index of any open model at 60 on Artificial Analysis, seven points clear of GLM-5.2, and on Arena it is still the highest-placed open entry, at #3 on WebDev behind Qwen3.8-Max-0902 and Claude Opus 5, and #3 on the Agent board. One catch decides which of the two you should actually use. The K3 download is 96 safetensors shards and about 1.56 TB, which needs a multi-node GPU cluster rather than a workstation, and it ships under a custom Kimi K3 License rather than MIT. So the practical recommendation for teams running their own weights is now GLM 5.3 Flash (Z.ai, MIT), released on August 26 at Intelligence Index 57, four points above GLM-5.2 and at less than half the size, 320B total with 18B active. It is the first natively multimodal model in the GLM-5 line, it takes a 1M-token context, and the weights were on Hugging Face under MIT the day it launched. GLM-5.2 (MIT) is the fallback, at Intelligence Index 53, and it still holds the independent record the newer model has not had time to earn, #6 on Arena WebDev (1,587.1) and #5 on the Agent board. Third place changed hands on the last day of July: DeepSeek V4-Flash 0731 was re-post-trained and jumped from Intelligence Index 40 to 52 on the current v4.1.1 index under MIT, though its API price has since gone up roughly fourfold, to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak. Three releases then closed the month, and none of them has an independent rating yet. GLM 5.3 finally published its weights on August 28, on exactly the date its placeholder had carried, but under a bespoke GLM 5.3 License rather than MIT, with a clause requiring any model-as-a-service operator above $10 billion in revenue to pass a Z.ai security review first; note also that GLM 5.3 Flash is not a trimmed GLM 5.3 but a separate model on a newly trained base, which is why it could ship under plain MIT while the flagship could not. Tencent open-sourced Hy4 Preview on August 28 at 770B total and 49B active under Apache 2.0, the largest permissively licensed model published so far, on the strength of an internal blind evaluation rather than any independent board. And Alibaba published Qwen3.8-Flash-Next on August 26, a 125B mixture-of-experts with only 6B active that previews the Qwen4 architecture, under the Qwen Community License 1.0 rather than Apache. All three are listed below rather than ranked.
| Model | Best For | Key Benchmark | Context / License | Where To Run |
|---|---|---|---|---|
| Kimi K3 | Highest-scoring open model, agentic and web work | II 60, highest of any open model; #3 Arena WebDev, #3 Arena Agent; 2.8T/104B active | 1M / Kimi K3 License | Hugging Face (96 shards, 1.56 TB), Moonshot API, providers |
| GLM 5.3 Flash | Best open model most teams can actually run | II 57 (v4.1.1), on the intelligence-versus-cost Pareto frontier at about $0.09 an index task; 320B/18B active, natively multimodal | 1M / MIT | Hugging Face, Z.ai API ($0.15/$0.50), OpenRouter |
| GLM-5.2 | Fallback with an independent track record | II 53; #6 Arena WebDev (1,587.1), #5 Arena Agent; 744B/40B active | 1M / MIT | Z.ai, Hugging Face, OpenRouter |
| GLM 5.3 | Open since August 28, but on a bespoke licence | Same 744B base as GLM-5.2, post-trained only, published checkpoint counted at 753B; CyberGym 84.5% and Terminal-Bench 3.0 28.3 (vendor figures) | 1M / GLM 5.3 License (MaaS above $10B revenue needs a Z.ai security review) | Hugging Face, Z.ai API ($1.40/$4.40), GLM Coding Plan, ZCode |
| DeepSeek V4-Flash 0731 | Best open value if you host it yourself | II 52 (v4.1.1), up from 40 on July 31; 284B/13B active | 1M / MIT | DeepSeek API ($0.22/$0.66 off-peak, $0.44/$1.32 peak since August 16), local |
| Tencent Hy4 Preview | Largest permissively licensed model yet | 770B/49B active; 2.99/4.00 on Tencent's own 203-task expert panel vs Kimi K3 2.94 and GLM 5.3 2.92 (vendor); no independent rating yet | 1M+ / Apache 2.0 | Hugging Face (BF16 and FP8), Tencent Cloud TokenHub, OpenRouter |
| Qwen3.8-Flash-Next | Preview of the Qwen4 architecture | 125B total plus a 51B N-gram embedding, 6B active; 91.7 GPQA Diamond and 62.5 SWE-Bench Pro (vendor); no independent rating yet | 262K to 1M / Qwen Community License 1.0 | Hugging Face, providers, self-host |
| LongCat-2.0 | Frontier open coder trained on Chinese chips | II 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active | 1M / MIT | Hugging Face, GitHub, OpenRouter |
| MiniMax M3 | Cheap frontier-class multimodal | II 44 (v4.1), 59% SWE-Bench Pro, multimodal | 1M / license TBD | Hugging Face, API $0.30/1M (50% off) |
| Nex-N2-Pro | Strongest open coding score | II 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B active | Qwen-based / Apache 2.0 | Hugging Face, providers, self-host |
| Kimi K2.7 Code | Strongest commercially-licensed open coder | +21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active | 256K / Modified MIT | Hugging Face, DeepInfra, providers |
| DeepSeek V4-Pro | Agentic real-world work | II 44 (v4.1), 1.6T/49B active | 1M / MIT | DeepSeek API ($0.66/$1.98 off-peak, $1.32/$3.96 peak since August 16), local |
| Hy3 | Newest permissive-licence entrant | II 41 (v4.1); #22 on Arena WebDev (1,516.4) | Apache 2.0 | Hugging Face, providers, self-host |
| Inkling | Thinking Machines' first open model | II 41 (v4.1), agentic 32.3, released July 15 | Open weights | Hugging Face, providers, self-host |
| Inkling Small | Same family at a quarter the size | II 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio in | Apache 2.0 | Hugging Face (BF16 and NVFP4), providers, self-host |
| NVIDIA Nemotron 3 Ultra | NVIDIA-tuned, fully permissive license | II 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active | 1M / OpenMDW | OpenRouter, Hugging Face, AWS (8x B200 self-host) |
| Qwen 3.5 (397B / 17B active) | Multimodal, fast decode | 88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v6 | 1M / open | Together, OpenRouter, local |
| Qwen3.6-35B-A3B | Efficient open agentic coder (3B active) | 86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active | 262K (to 1M YaRN) / Apache 2.0 | Hugging Face, OpenRouter, local |
| Qwen3.6-27B | Laptop-runnable dense coder | 87.8 GPQA Diamond, dense 27B, multimodal | 256K / Apache 2.0 | Local Mac/PC, Hugging Face, OpenRouter |
| Rio 3.5 Open 397B | Qwen 3.5 fine-tune, multilingual reasoning | 70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5 | 397B/17B active / MIT | Hugging Face, providers, self-host |
| Llama 4 Maverick | Meta-line flagship | 17B active / 400B total params | Llama 4 license | Meta cloud, Hugging Face, local |
| NVIDIA Nemotron 3 Nano Omni | Edge / low-power | Multimodal, very small footprint | Compact / open | Local, NVIDIA tool |
Licensing matters as much as raw score here, and August widened the gap between the two groups. Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM 5.3 Flash (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Hy4 Preview (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use with no revenue test. The rest carry conditions worth reading before you build on them: MiniMax M3 and MiniMax H3 ship under their own community licences, H3 additionally requiring an application from the US, EU, UK and South Korea; Kimi K3 sits on a custom licence with a revenue threshold for anyone reselling it as a service; GLM 5.3 requires a Z.ai security review of any model-as-a-service operator above $10 billion in revenue; and the Qwen Community License 1.0 on Qwen3.8-Flash-Next requires a separate licence from Qwen to run a model-as-a-service or an AI work-assistant business, though internal use is exempt. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.
Benchmarks, Prices, and Hands-On Use
Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 2.1 board, covering the Intelligence and Agentic indexes, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.
Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We rank on measured scores: a category crown moves as soon as a model out-scores the holder on the boards that define that category, without waiting for the vote-based boards to catch up, and a model that is not yet widely available is flagged on its card rather than denied one. We re-check ranks as well as figures, because a score can stay correct while the board around it moves. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.
Fello AI brings the top AI models together in one lightweight native app.
Switch between them anytime, no separate subscriptions.
Download Fello AI




