Best AI Models in 2026
Rankings, comparisons, and deep dives, updated monthly as new models ship.
Google announced Gemini 4 Argon on September 30, its first flagship since Gemini 3.1 Pro, and the independent boards put it straight into the top tier: Artificial Analysis scores it 52.6 on Intelligence Index v4.3.2, level with GPT-6 Astra and behind only Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1, with the lowest hallucination rate it has measured among leading models, and Arena ranks it first on its text and creative-writing boards. It wins no category on this page, for two reasons: most of the boards our categories are decided on have not rated it yet, and only vetted cyber defenders in Google's Fairwind Program can use it. Google says paid API customers and Google AI Ultra subscribers come next, with no date, at an introductory $2 / $10 per 1M tokens.
Anthropic still holds most of the measured crowns. Claude Opus 5.5 (September 22) is #1 on the Intelligence Index at 57.6 and on Arena's WebDev board, and keeps the coding crown at $4 / $20; Claude Sonnet 5.5 (September 28) is second at 56.0, leads Terminal-Bench 4.0 at the same $2 / $10 as Sonnet 5, and is the Claude model you can use free on claude.ai. Claude Fable 5.1 keeps writing and accuracy on the boards those categories are decided on. This page ranks on measured scores, so a crown moves as soon as a model out-scores the holder, without waiting for the vote-based boards to catch up. Our guide to AI benchmarks explains how that index is built and why its version number matters.
OpenAI's answer is price. GPT-6 Sol and GPT-6 Luna (September 22) halved its API rates, and GPT-6.1 Sol (September 29) scores 51.8 at the same $2 / $10, sixth among distinct models but the cheapest in the top tier at $0.72 an index task. ARC Prize now puts it second on ARC-AGI-2 at 94.2%, 0.8 points behind GPT-6 Astra, and Arena third on WebDev. None of the GPT-6 models is in ChatGPT's chat window yet: they run in ChatGPT Work and Codex.
SpaceXAI's Grok 4.7 (September 21) reached the API, Cursor and Grok Build at an unchanged $2 / $6 while the Grok app still serves 4.6, and Xiaomi's MiMo-V2.6-Pro (September 21, plain MIT) is still the highest-scoring open model at 46.3 and $0.13 an index task. Below are the category winners for October 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks and pricing are updated within 48 hours of any major model launch.
Want the top AI models without juggling separate subscriptions? Fello AI brings the leading models together in one native app for Mac, iPhone and iPad.
Download Fello AISep 30 Google Gemini 4 Argon tops Arena's text board and matches GPT-6 Astra on the Intelligence Index, but only cyber defenders can use it New
Sep 29 OpenAI GPT-6.1 Sol reaches fifth on the Intelligence Index at $2 / $10, the cheapest per task in the top tier New
Sep 28 Anthropic Claude Sonnet 5.5 lands second on the Intelligence Index at 56.0, at Sonnet's $2 / $10 New
Sep 22 OpenAI GPT-6 Sol and GPT-6 Luna halve OpenAI's API prices, and land everywhere except ChatGPT's chat window New
Sep 22 Anthropic Claude Opus 5.5 takes the Intelligence Index at 57.6 and undercuts Opus 5 on price New
Sep 21 SpaceXAI Grok 4.7 arrives on a larger base model at an unchanged $2 / $6, and costs 47% more per task New
Sep 21 Xiaomi Xiaomi's MiMo-V2.6-Pro is the highest-scoring open model ever measured, at $0.13 an index task under MIT New
Sep 15 TypeSafe TypeSafe leaves stealth with Jev, a model that returns typed decisions instead of text, at $0.042 per 1M tokens New
Sep 10 DeepSeek DeepSeek retires V4-Flash, replaces it with V4.1-Flash and cuts the Flash rate by about a third New
Sep 3 OpenAI GPT-6 Astra ships at $10 / $50, and within a day it is on Plus and Pro and leading the coding boards New
Sep 2 Alibaba Qwen3.8-Max-0902 takes Arena's WebDev board off Claude Opus 5 by three points, at $2 / $6 New
Sep 2 Muse Spark 1.3, Intelligence Index 57 to 61 at an unchanged $1.25 / $4.25, winning coding and losing every agent row New
Sep 2 Google Gemini 3.8 Flash, three points on the Intelligence Index at the same $0.75 / $3.75 New
Sep 1 Anthropic Claude Fable 5.1 and Mythos 5.1, a new #1 on Artificial Analysis at 66 and a 75% cut to cache reads New
Pending Google Gemini 3.5 Pro, still unreleased, and Gemini 4 Argon has overtaken it Delayed
Best AI for Writing
The best AI for writing is Claude Fable 5.1, and the case for it is consistency rather than a single first place. It is inside the top ten of all three prose boards: second on EQ-Bench Creative Writing v3 at 2162.0, third on LiveBench Language at 89.5 and tenth on Arena's creative-writing board. No other model places that well across all three. Gemini 4 Argon now leads Arena's creative-writing and text boards, at 1521.8 and 1524.8, but neither EQ-Bench nor LiveBench has rated it and only Google's Fairwind partners can use it. Claude Opus 5.5 is second on Arena's creative-writing board at 1515 but seventh on EQ-Bench and eighth on LiveBench Language; GPT-6 Astra leads EQ-Bench at 2173.3 but sits 39th on Arena's creative-writing board; Claude Fable 5 tops LiveBench Language at 90.7 and is third on Arena's creative-writing board, but eleventh on EQ-Bench. If your writing is a work deliverable rather than prose, the GDPval-AA v2.1 professional-deliverables board is a different question, and Claude Opus 5.5 leads it at 1846, with Claude Sonnet 5.5 two points behind at 1844 and Claude Fable 5.1 at 1735. Claude Sonnet 5.5 is the free-tier value pick, since anyone can use it on claude.ai, though of the three prose boards only LiveBench Language has rated it so far, at 83.4. Claude Fable 5 is the pick if you weight Arena's human votes above the harness scores; it costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. Translation is a separate question with a separate winner, which we work through in our guide to the best AI for translation.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Fable 5.1 | Best writing overall | #2 EQ-Bench Creative Writing (2162.0), #3 LiveBench Language (89.5) and #10 Arena creative writing, the best combined placing of any model | Leads none of the three outright; Gemini 4 Argon now leads Arena's creative-writing board, with Claude Opus 5.5 second | $10 / $50 |
| Claude Fable 5 | Strong on the human-vote boards | #1 LiveBench Language (90.7), #3 Arena creative writing (1503.0), #3 Arena text (1504.9) | #11 EQ-Bench; priciest option here | $10 / $50 |
| Kimi K3 | Creative fiction and voice | #5 EQ-Bench Creative Writing (2082.3) | Only #26 on Arena creative writing | $3 / $15 |
| Claude Sonnet 5.5 | Free everyday writing | Free on claude.ai, 1M context; #2 GDPval-AA v2.1 (1844), two points behind Opus 5.5 | 83.4 LiveBench Language; not yet rated on EQ-Bench or Arena | $2 / $10 |
| Claude Opus 5.5 | Arena's top creative writer you can use | #2 Arena creative writing (1515) and #4 Arena text (1504); #1 GDPval-AA v2.1 (1846) | #7 EQ-Bench and #8 LiveBench Language | $4 / $20 |
| Gemini 4 Argon | Not available yet | #1 Arena creative writing (1521.8) and #1 Arena text (1524.8) | Unrated on EQ-Bench and LiveBench; Fairwind partners only | $2 / $10 intro, $4 / $20 standard |
| GPT-5.5 | Fact-anchored business writing | Documented factual-reliability gains over GPT-5.4 | Reasoning tiers now marked deprecated | $5 / $30 |
| Gemini 3.8 Flash | Bulk drafts at scale | Intelligence Index 40.9, 308.4 tok/s output, 1M context | Agent-tuned rather than prose-tuned; intro price doubles January 1, 2027 | $0.75 / $3.75 |
The writing crown stays with Claude Fable 5.1, but its margin on Arena narrowed. Google's Gemini 4 Argon, announced September 30, went straight to first on Arena's creative-writing board at 1521.8, pushing Claude Opus 5.5 to second and Claude Fable 5 to third, and Fable 5.1 slipped from seventh to tenth, so this card's claim is now the top ten of all three prose boards rather than the top seven. Argon cannot take the crown: EQ-Bench and LiveBench have not rated it, and only Google's Fairwind partners can use it. In September, Claude Opus 5.5 and Claude Sonnet 5.5 arrived, Sonnet 5.5 replaced Sonnet 5 as the free pick, and GPT-6.1 Sol entered LiveBench Language second, pushing Fable 5.1 to third there.
Best AI for Chat & Daily Assistant
The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #20 on Arena's text leaderboard at 1483.8, where Gemini 4 Argon now leads at 1524.8 and Claude Opus 5.5 is fourth at 1504.0. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (the 5.6 tiers: Luna $0.20 / $1.20, Terra $2 / $12, Sol $4 / $20 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| GPT-5.6 | Everyday chat, ChatGPT's default | The assistant most people can open; Terra matches GPT-5.5 at ~half cost | #20 on Arena text; scheming flagged by METR | Free / $20/mo Plus; API $0.20 / $1.20 to $4 / $20 |
| Claude Opus 5.5 | Highest-scoring conversation | #1 on the Intelligence Index (57.6) and #4 on Arena text (1504.0) | Not on the Free plan; usage-limited on Pro | $20/mo Pro, $4 / $20 API |
| Gemini 4 Argon | Not available yet | #1 on Arena text (1524.8); 15% hallucination rate, the lowest among leading models | Fairwind partners only; not in the Gemini app | $2 / $10 API intro, once released |
| Claude Sonnet 5.5 | Free Claude chat | Free on claude.ai; #2 on the Intelligence Index (56.0), behind only Claude Opus 5.5 | Not yet rated on Arena text; runs at Medium effort by default in the apps | Free / $20/mo Pro, $2 / $10 API |
| GPT-5.5 Instant | Hallucination-sensitive daily work | 52.5% fewer hallucinated claims vs 5.3 Instant | Reasoning tiers now marked deprecated | $20/mo Plus; API $5 / $30 |
| Gemini 3.6 Flash | Fast, free, multimodal | Still the model the free Gemini tier serves, 1M context, #21 on Arena text | Weaker on hardest reasoning; 3.8 Flash outscores it at the same price | Free / $0.75 / $3.75 API |
| Fello AI | All the top models, one app | ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad | Routed via app, not direct | $9.99/mo |
GPT-5.6 keeps the chat pick on reach, and nothing about who can open what changed: GPT-5.6 Luna is still ChatGPT's free default, and the GPT-6 models still run in ChatGPT Work and Codex rather than in Chat. The preference board moved again. Gemini 4 Argon, announced September 30, took first on Arena's text board at 1524.8, and Claude Opus 5.5, which led it a week ago, is now fourth at 1504.0, about a point behind Claude Opus 4.6 and Claude Fable 5. Opus 5.5 keeps the second podium place as the top scorer on Artificial Analysis's index; Argon stays off the podium because only Google's Fairwind partners can use it and it is not in the Gemini app. In September, Claude Sonnet 5.5 became the free Claude model and took the third podium place from Claude Opus 5.
Best AI for Images
The best AI for image generation is ChatGPT Images 2.5, which OpenAI made the default on every ChatGPT tier on September 8 and which leads all four image boards we track. Its two models finish first and second on each: GPT-Image-2.5 Sunburst, built for precision, leads Arena text-to-image at 1423.6 and Arena image editing at 1522.4, and it leads Artificial Analysis's quality Elo at 1196.5 and its image-editing board at 1221.3, with the faster GPT-Image-2.5 Flare second on all four. When it launched, Artificial Analysis had rated neither model; it has since rated both, and its boards now agree with Arena's. Read the Sunburst-to-Flare gap on text-to-image as provisional, because each rests on about 11,000 votes against GPT Image 2's 89,000. MAI-Image-2.6 is the runner-up, third on Artificial Analysis's image-editing board at 1191.6 and fourth on Arena's text-to-image board at 1335.0, with Meta's Muse Image third. Reve 2.1 keeps its case for layout, typography and native 4K but has left the podium: it is fifth on Arena text-to-image and fourteenth on Arena image editing. Our full breakdown of what shipped is in ChatGPT Images 2.5.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| ChatGPT Images 2.5 | Images with readable text | #1 and #2 on all four boards: Arena text-to-image (1423.6 / 1400.9), Arena editing (1522.4 / 1478.5), Artificial Analysis quality (1196.5 / 1189.8) and editing (1221.3 / 1207.0) | Its Arena text-to-image lead rests on about 11,000 votes per model | Default on every ChatGPT tier |
| MAI-Image-2.6 | Editing an image you already have | #3 on Artificial Analysis image editing (1191.6) and #4 on Arena text-to-image (1335.0) | Public preview, not in Copilot or Bing yet; lost the Artificial Analysis editing lead to ChatGPT Images 2.5 | MAI Playground / Microsoft Foundry |
| Muse Image | Image editing, Meta ecosystem | #5 on Artificial Analysis image editing (1171.3) and #7 on its quality Elo | New, thin tooling around it | Meta AI app |
| Nano Banana 2 (Gemini 3.1 Flash Image) | Photoreal portraits and products | #4 on Artificial Analysis text-to-image, ahead of Nano Banana Pro on Arena's | Nano Banana Pro is ahead on Arena image editing | Gemini app / AI Studio |
| Seedream 5.0 Pro | Multilingual text + region editing | 10+ languages incl. Arabic RTL, lasso and layer editing | Around tenth on the independent boards; copyright cloud | BytePlus / Magnific |
| Midjourney v8 | Stylized art, illustration | Aesthetic baseline most artists prefer | Weaker on text in image | $10-$120/mo |
| Grok Imagine | NSFW / Spicy Mode | Most permissive guardrails; Arena rates its Image 2.0 model #6 on text-to-image and #4 on image editing, and Artificial Analysis now has it #4 on quality Elo | Nothing in the release record shows Image 2.0 shipping as a product | $30/mo SuperGrok |
ChatGPT Images 2.5 keeps the image crown and still holds first and second on all four boards we track. Since the last update the Arena figures have drifted by a few points either way, Sunburst now reading 1423.6 on text-to-image and 1522.4 on editing, and Reve 2.1 and Grok Imagine Image 2.0 swapped fifth and sixth on Arena text-to-image. Gemini 4 Argon is a text model and changes nothing here. In September, Images 2.5 replaced GPT Image 2 at the top on September 8, and Artificial Analysis rated both of its models, which ended the split between its boards and Arena's.
Best AI for Video
The best AI for video generation is Gemini Omni 1.1 Flash, which Google shipped on August 27 and which leads Arena's text-to-video board at 1516, just ahead of its own predecessor Gemini Omni Flash at 1513. Read that as a tie: the two sit inside each other's confidence intervals and Arena itself gives 1.1 Flash a rank band of one to four. What breaks it is specification, where 1.1 Flash is clearly the more capable model: scene extension in 10-second increments to a cumulative 40 seconds, first and last frame control, video references, 360p drafts and 1080p or 4K output, priced at roughly $0.03 per second at 360p, $0.10 at 720p, $0.15 at 1080p and $0.30 at 4K. Gemini Omni Flash remains the cheaper option at about $0.10 a second but caps at 10-second clips. Two limits worth knowing. Artificial Analysis has not rated 1.1 Flash at all; on its text-to-video board the older Omni Flash leads at 1512.8 from Alibaba's Wan 3.0 at 1431.4 and MiniMax H3 at 1407.7. And on Arena's image-to-video board 1.1 Flash is second at 1488.5 behind MiniMax H3 at 1495.4, so the crown rests on text-to-video and on specification, not on a sweep. If you need longer takes inside Google's tooling, Veo 3.1 still runs in the Gemini app, AI Studio and Vertex AI with native audio and 1080p output. MiniMax H3 (July 31) remains the specification contender, with 2K clips of 4 to 15 seconds and native stereo audio from $0.13 a second.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Gemini Omni 1.1 Flash | Best AI video overall | #1 Arena text-to-video (1516); 40-second scene extension, 4K output | Unrated by Artificial Analysis; second to MiniMax H3 on Arena image-to-video | $0.03-$0.30/sec; Gemini API / AI Studio / Flow |
| Gemini Omni Flash | Cheaper 10-second clips | #2 Arena text-to-video (1513), #2 AA video (1238); conversational editing | Caps at 10-second generations | ~$0.10/sec; Gemini app / AI Studio |
| Wan 3.0 | Video from documents and slides | #1 on Artificial Analysis's video board (1239); 2-30s, up to 1080p, Omni-Reference input | API-only, no published weights; #6 on Arena | $0.05-$0.20/sec |
| MiniMax H3 | 2K clips with native audio | 4-15s at 2K, native stereo audio; #8 Arena text-to-video (1460) | #4 on AA video (1227); the open weights exclude the 2K upscaler and need an application in the US, EU, UK and South Korea | $0.13/sec at 2K |
| Dreamina Seedance 2.5 | ByteDance challenger | #7 Arena text-to-video (1474); leads image-to-video with audio | ByteDance ecosystem, limited Western access | Dreamina / BytePlus |
| Muse Video | Meta ecosystem video | #9 on Arena text-to-video at 1456 | Newest of the group, thin tooling | Meta AI app |
| Veo 3.1 | Longer production clips | Native audio, 1080p, strong physics consistency | #12 on Arena video, not the quality leader | Google AI Pro / Ultra |
| Kling 3.0 / 3.0 Turbo | Fast iteration at lower cost | Native 4K, 60fps, 15-second clips; Turbo shipped June 17 | Outside the top 16 on Arena text-to-video | From $10/mo |
| Luma Ray 3 | Photoreal scenes | Strong realism for landscapes | Smaller community | Free / from $9.99/mo |
Gemini Omni 1.1 Flash keeps the video crown at 1516 on Arena text-to-video, still a statistical tie with Gemini Omni Flash at 1513 broken on capability. Two models that shipped months ago have now gathered enough Arena votes to rank, on about 1,200 to 1,300 votes each: Black Forest Labs' FLUX 3 Video enters third at 1493 and SpaceXAI's Grok Imagine Video 1.5 fourth at 1492, which pushes Wan 3.0 to sixth and Meta's Muse Video from third to ninth. MiniMax H3 still leads Arena's image-to-video board at 1495.4, ahead of Omni 1.1 Flash at 1488.5. In September, Artificial Analysis re-anchored its video Elo, and it has still not rated Omni 1.1 Flash.
Best AI for Coding
The best AI for coding is Claude Opus 5.5, which Anthropic shipped on September 22, though it no longer leads every board, and the model that took two of them is its own cheaper sibling. Claude Sonnet 5.5, released September 28, leads Artificial Analysis's own Terminal-Bench 4.0 run at 63.6% against Opus 5.5's 59.6% and LiveBench Coding at 91.4 against 89.3. Opus 5.5 keeps the crown on the rest of the evidence. It leads Arena's WebDev board at 1818, ahead of GPT-6 Astra at 1789 and GPT-6.1 Sol at 1759. It scores 71.7 on LiveBench Agentic Coding, second only to DeepSeek V4.1-Flash, where Sonnet 5.5 manages 56.3, and on Anthropic's own harnesses it beats Sonnet 5.5 on FrontierCode v1.1, 54.4% to 52.1%, and on CursorBench 4.0, 57.8% to 55.5%. Price does not settle it the way the token rates suggest. Sonnet 5.5 is $2 / $10 per 1M tokens against Opus 5.5's $4 / $20, but at max effort it spends so many tokens that an Intelligence Index task costs $7.60 against Opus 5.5's $5.98. Its value is at lower effort: at xhigh it scores 51.9 on the index for $2.74 a task and at high 46.7 for $1.08. GPT-6 Astra keeps Arena's image-to-WebDev board at 1733 and matches Opus 5.5's 59.6% on Terminal-Bench 4.0 at xhigh, at $10 / $50, and Claude Fable 5.1 is fourth on WebDev and second on image-to-WebDev. GPT-6.1 Sol is the cheap OpenAI option: third on WebDev and 56.1% on Terminal-Bench 4.0 for $0.72 an index task, though LiveBench puts it well outside the top ten on Coding, at 80.7. Gemini 4 Argon leads Google's own DeepSWE v1.1 table at 77.9%, but Artificial Analysis measures 57.1% on Terminal-Bench 4.0, Arena puts it eighth on WebDev, and it is not available outside Google's Fairwind program. Among open weights, Xiaomi's MiMo-V2.6-Pro takes 34.8% on Terminal-Bench 4.0 under plain MIT at $0.13 an index task, GLM-5.3 is still the strongest open model on that harness at 41.9%, and Kimi K3 is the highest-placed open model on Arena WebDev, twelfth at 1658, on only 12.6% of the harness. Artificial Analysis has retired both its Coding Index and its Coding Agent Index, so the two composite coding rankings this page used to quote no longer exist.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Opus 5.5 | Best coding overall | #1 Arena WebDev (1818) and #1 Intelligence Index (57.6, v4.3.2); 71.7 LiveBench Agentic Coding, 15 points above Sonnet 5.5 | Second to Sonnet 5.5 on Terminal-Bench 4.0 (59.6% to 63.6%) and LiveBench Coding (89.3 to 91.4) | $4 / $20 |
| Claude Sonnet 5.5 | Value pick for coding | #1 Terminal-Bench 4.0 (63.6%) and #1 LiveBench Coding (91.4); 51.9 on the index for $2.74 a task at xhigh | 56.3 LiveBench Agentic Coding; at max effort $7.60 an index task, more than Opus 5.5; only #5 on Arena WebDev (1709) | $2 / $10 |
| GPT-6 Astra | Top of Arena image-to-WebDev | #1 Arena image-to-WebDev (1733), #2 WebDev (1789); 59.6% Terminal-Bench 4.0 at xhigh | 2.5x Opus 5.5's price; behind Sonnet 5.5 on Terminal-Bench and well behind both Claude models on LiveBench Coding | $10 / $50 |
| Claude Fable 5.1 | Highest composite score before Opus 5.5 | #4 Arena WebDev and #2 image-to-WebDev; Intelligence Index 53.4; AA-Briefcase 1678 | 52.0% Terminal-Bench 4.0; 2.5x the price of Opus 5.5 | $10 / $50 |
| Claude Opus 5 | Superseded by Opus 5.5 | 49.0% Terminal-Bench 4.0 and #6 Arena WebDev and #3 image-to-WebDev | Costs more per token than Opus 5.5 and scores seven points lower | $5 / $25 |
| Claude Fable 5 | Hardest long-horizon agentic work | Intelligence Index 49.6; 42.4% Terminal-Bench 4.0; 1M context | The priciest per index task on this table at $8.75; #6 on Arena image-to-WebDev | $10 / $50 |
| Qwen3.8-Max-0902 | Brief holder of Arena's WebDev board | #10 Code Arena: WebDev (1670); Intelligence Index 45.4; 2.4T parameters, 1M context | QwenCloud API only with no consumer front-end; lost the WebDev lead within days | $2 / $6 |
| Kimi K3 | Highest-placed open model on Arena | #12 Arena WebDev (1658), the best open placement on that board; 2.8T/104B active | Only 12.6% on Terminal-Bench 4.0; 1.56 TB to self-host | $3 / $15 |
| GPT-5.6 Sol | OpenAI's cheaper flagship | Intelligence Index 47.0; 39.9% Terminal-Bench 4.0 | Superseded by GPT-6 Sol, which scores higher at half the price; eval-gaming flagged by METR | $4 / $20 |
| GPT-6 Sol | Cheap GPT-6 coding in Codex | Intelligence Index 47.5 (v4.3.2) and 43.9% Terminal-Bench 4.0, both above GPT-5.6 Sol, at half its price | #7 Arena WebDev (1689); in Codex and ChatGPT Work only, not in Chat | $2 / $10 |
| GPT-6.1 Sol | Cheapest top-tier model per task | Intelligence Index 51.8 (v4.3.2), #3 Arena WebDev (1759) and 56.1% Terminal-Bench 4.0 at $0.72 an index task, with cached input at $0.10 | 80.7 on LiveBench Coding, well outside the top ten; in Codex and ChatGPT Work only, not in Chat | $2 / $10 |
| Gemini 4 Argon | Not available yet | 77.9% DeepSWE v1.1 on Google's own figures; Intelligence Index 52.6 at $1.99 an index task | 57.1% Terminal-Bench 4.0 and #8 Arena WebDev (1679); Fairwind partners only | $2 / $10 intro, $4 / $20 standard |
| Muse Spark 1.3 | Cheap agentic coding | Intelligence Index 48.1 at $1.25 / $4.25, and $1.60 per index task | 33.3% on Terminal-Bench 4.0; the coding wins come from Meta's own chart, measured on a partners-only max variant | $1.25 / $4.25 |
| Grok 4.7 | Cheap value coder | 46.3% CursorBench 4.0 and 71.0% DeepSWE v1.1 on SpaceXAI's own figures; Intelligence Index 46.3 | 24.7% on Terminal-Bench 4.0; prompts of 200K tokens and up re-bill the whole request at double | $2 / $6 |
| Gemini 3.8 Flash | Agent coding at scale | Intelligence Index 40.9; 308.4 output tokens per second | 19.7% on Terminal-Bench 4.0; intro price doubles January 1, 2027 | $0.75 / $3.75 |
| GLM 5.3 Flash | Cheapest serious coder you can host | 32.8% Terminal-Bench 4.0 at $0.25 an index task; MIT, 320B/18B active | GLM-5.3 reaches 41.9% on the same harness; no consumer product | Open weights (MIT) |
| MiMo-V2.6-Pro | Highest open Intelligence Index | 46.3 on v4.3.2 and 34.8% Terminal-Bench 4.0 at $0.13 an index task; MIT, 1.02T/42B active | Only #23 on Arena WebDev (1619), eleven places behind Kimi K3 | $0.435 / $0.87 |
Claude Opus 5.5 keeps the coding crown. Since the last update the move is on Arena: GPT-6.1 Sol entered WebDev third at 1759, ahead of Claude Fable 5.1 and Claude Sonnet 5.5, and Gemini 4 Argon entered eighth at 1679, which pushed Qwen3.8-Max-0902 to tenth and Kimi K3 to twelfth. Argon leads Google's own DeepSWE v1.1 table at 77.9%, but on Artificial Analysis's Terminal-Bench 4.0 it scores 57.1%, behind Sonnet 5.5, Opus 5.5 and GPT-6 Astra, and nobody outside Google's Fairwind program can use it. In September the crown moved twice, from Claude Fable 5.1 to GPT-6 Astra and then to Opus 5.5, and Claude Sonnet 5.5 took Terminal-Bench 4.0 and LiveBench Coding as the value pick.
Best AI for Creativity
The best AI for unfiltered, on-trend creative work is Grok 4.7, and we want to be exact about why. This pick is about the product line, not the prose quality. Grok carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. Read the availability carefully: SpaceXAI shipped Grok 4.7 on September 21 to the Grok API, Cursor, Grok Build, third-party harnesses and cloud routers, and its announcement says nothing about the consumer app, which still serves Grok 4.6 to SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, though the picture is more mixed than it was. EQ-Bench has now rated Grok 4.7 and puts it eighth at 2006.7, inside its top ten and far above Grok 4.5 in 52nd, but Arena's human voters put 4.7 only 68th on creative writing, behind Grok 4.6 in 49th and the older Grok 4.20-beta1 in 24th. If you are picking on output quality alone, Gemini 4 Argon now leads Arena's creative-writing board, though almost nobody can use it yet, with Claude Opus 5.5 second and Claude Fable 5 third, while EQ-Bench puts GPT-6 Astra first at 2173.3, Claude Fable 5.1 second and Kimi K3 fifth. Choose Grok for what it will let you make, not for how well it writes.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Grok 4.7 | Unfiltered, opinionated, on-trend | Fewest content restrictions, native real-time X grounding | #68 on Arena creative writing, though #8 on EQ-Bench (2006.7) | $2 / $6 API; $30/mo SuperGrok app, still on 4.6 |
| Claude Fable 5 | Highest-quality creative prose | #3 Arena creative writing, #1 LiveBench Language | Cautious guardrails on edgy briefs | $10 / $50 API |
| Kimi K3 | Fiction and distinctive voice | #5 EQ-Bench Creative Writing at 2082.3 | Only #26 on Arena creative writing | $3 / $15 |
| Claude Opus 5 | Long-form structured creativity | Holds long threads and self-edits; Intelligence Index 50.8, now behind six newer models | Most cautious of the group | $20/mo Pro; $5 / $25 API |
| Gemini 3.1 Pro | Multimodal creative | Strong text, image and video chain | Quotas inside the Gemini app | Free / $2.00-$4.00 API in |
| Grok Imagine (Spicy Mode) | NSFW / adult creative | Most permissive image generation | Niche use case | $30/mo SuperGrok |
Grok keeps this pick on Grok 4.7, and nothing about the product's permissiveness or its live X access changed, which is what the pick rests on. Arena's creative-writing board moved around it: Gemini 4 Argon, announced September 30, went in at first, Claude Opus 5.5 is now second and Claude Fable 5 third, and Grok 4.7 edged up from 72nd to 68th. EQ-Bench has not changed and still puts 4.7 eighth. The split surface still matters for this category: 4.7 is on the API, Cursor and Grok Build, while the $30/month Grok app that the permissiveness argument is really about still serves 4.6. In September, Grok 4.7 replaced Grok 4.6 as SpaceXAI's flagship on September 21 and both creative-writing boards rated it.
Best AI for Accuracy & Research
The best AI for accuracy and research is Claude Fable 5.1, which holds the highest factual accuracy Artificial Analysis has measured, 67% on AA-Omniscience. That is the exact column this category is decided on. Claude Opus 5.5 beat Fable 5.1 on almost everything else when it arrived on September 22, including Humanity's Last Exam at 61.4% to 59.1% and the hallucination-penalised AA-Omniscience index at 46.4 to 43.5, but on raw factual accuracy it reads 66% to Fable 5.1's 67%. Treat that one-point gap as a tie and read the two together: Opus 5.5 is the better bet when a wrong answer is worse than no answer, Fable 5.1 when you want the most facts right. The value pick is now GPT-6.1 Sol, which answers 62% correctly with an AA-Omniscience index of 41.5 at $2 / $10 per 1M tokens and $0.72 an index task, and scores 98.5% on ARC Prize's ARC-AGI-1 for $0.06 a task. That beats Gemini 3.1 Pro, our value pick until this update, on every accuracy measure Artificial Analysis publishes: it reads 55% raw accuracy and an index of 31.9. Gemini 3.1 Pro keeps its case where the answer has to be current, because it grounds answers in Google Search and is in the free Gemini app, while GPT-6.1 Sol runs in ChatGPT Work, Codex and the API rather than in ChatGPT's chat window. Gemini 4 Argon is the model to watch: Artificial Analysis measures a 15% hallucination rate, the lowest among leading models, because it declines rather than guesses, but it answers only 50% correctly and only Google's Fairwind partners can use it. On grounded search specifically, Arena's search leaderboard is led by OpenAI: GPT-5.6 Sol tops it at 1257, with Claude Fable 5 fifth and Gemini 3.1 Pro grounding ninth. On novel reasoning the crown sits elsewhere, with GPT-6 Astra leading ARC-AGI-2 and ARC-AGI-3. Claude Opus 5 holds no accuracy crown because Artificial Analysis measures its hallucination rate at 61%.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| Claude Fable 5.1 | Highest measured factual accuracy | 67% factual accuracy on AA-Omniscience, the highest Artificial Analysis has measured, one point above Claude Opus 5.5 | No cheap tier; unrated on Arena's search leaderboard; Opus 5.5 leads Humanity's Last Exam and the AA-Omniscience index | $10 / $50 |
| Claude Opus 5.5 | When a wrong answer is worse than none | #1 AA-Omniscience index (46.4) and #1 Humanity's Last Exam (61.4%); 66% raw accuracy | One point behind Fable 5.1 on raw accuracy | $4 / $20 |
| GPT-6.1 Sol | Best value | 62% raw accuracy and an AA-Omniscience index of 41.5 at $0.72 an index task; 98.5% ARC-AGI-1 at $0.06 a task | 54% hallucination rate; in ChatGPT Work and Codex, not in Chat | $2 / $10 |
| Gemini 4 Argon | Fewest hallucinations, once available | 15% hallucination rate on AA-Omniscience, the lowest among leading models | Only 50% raw accuracy; Fairwind partners only | $2 / $10 intro |
| Gemini 3.1 Pro | Grounded search, free in the Gemini app | Native Google Search grounding; 98% ARC-AGI-1 at $0.52/task | 55% raw accuracy on AA-Omniscience, below GPT-6.1 Sol; #9 on Arena search | $2.00-$4.00 / $12.00-$18.00 (tiered) |
| Claude Fable 5 | Grounded search | #5 on Arena's search leaderboard, behind GPT-5.6 Sol | No single cheap tier | $10 / $50 (Fable 5) |
| GPT-5.6 Sol | Novel reasoning | #4 ARC-AGI-2 at 92.5%, behind GPT-6 Astra (95%), GPT-6.1 Sol (94.2%) and Claude Opus 5.5 (93.3%) | Scheming flagged by METR | $4 / $20 |
| Claude Opus 5 | Hardest unseen problems | 30.2% ARC-AGI-3 on the standard harness, second to GPT-6 Astra (ARC Prize) | Artificial Analysis measures a 61% hallucination rate | $5 / $25 |
| Qwen 3.7 Max | Frontier accuracy at value pricing | 92.4 GPQA Diamond, 200 free requests/day | API-only, no chat front-end | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Opus 4.6 | Honesty under pressure | #1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5 | Superseded as a flagship | Legacy Anthropic model |
The accuracy crown stays with Claude Fable 5.1, at 67% raw factual accuracy against Claude Opus 5.5's 66%, and the value pick moved. GPT-6.1 Sol now beats Gemini 3.1 Pro on every accuracy measure Artificial Analysis publishes, 62% raw accuracy against 55% and an AA-Omniscience index of 41.5 against 31.9, and ARC Prize scores it 98.5% on ARC-AGI-1 for $0.06 a task against Gemini 3.1 Pro's 98% for $0.52, so it takes the value card. Gemini 3.1 Pro stays in the table for grounded search and the free Gemini app. Gemini 4 Argon, announced September 30, brings the lowest hallucination rate Artificial Analysis has measured among leading models, 15%, but it answers only half its questions correctly and is not yet available, so it changes nothing at the top. Two older lines are corrected as well: Claude Opus 5's hallucination rate now reads 61% rather than 50%, and it no longer leads ARC-AGI-3. In September, Opus 5.5 took the hallucination-penalised index and Humanity's Last Exam, which is why this block now splits the two Claude models rather than crowning one outright.
Best AI for Problem Solving
The best AI for hard problem solving is GPT-6 Astra, which took the reasoning boards off GPT-5.6 Sol within days of launching. It leads LiveBench Reasoning at 92.65 and ARC-AGI-2 at 95%, the highest anyone has scored against that board's 100% human panel, and sits fourth on LiveBench Mathematics at 96.81, behind Claude Opus 5.5 at 97.08, Claude Fable 5.1 at 97.01 and GPT-6.1 Sol at 96.83. It also reports 97.6% on FrontierMath Tier 4 v2, a number worth carrying with Epoch AI's disclosure that OpenAI funded the benchmark and holds exclusive access to part of it. The catch is price and reach: $10 / $50 per 1M tokens against GPT-6.1 Sol's $2 / $10, and it needs a paid ChatGPT plan. GPT-6.1 Sol, released September 29, is the value alternative and very nearly a co-leader: 92.63 on LiveBench Reasoning, 0.02 behind Astra, and 94.2% on ARC-AGI-2, 0.8 points behind, for $0.25 a task against Astra's $1.12. Astra keeps the crown because it still leads both boards and ARC-AGI-3, where it scores 62.7% on the standard harness against GPT-6.1 Sol's 52.7%. Gemini 4 Argon leads Arena's math category, but on only a few hundred votes, and neither LiveBench nor ARC Prize has rated it. GPT-5.6 Sol, the old value pick, is now fifth on LiveBench Reasoning and fourth on ARC-AGI-2 at 92.5%. Qwen 3.7 Max is the budget pick for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, with 200 free model requests a day. Claude Opus 5 remains the alternative for long agentic reasoning chains, though Astra has taken ARC-AGI-3 from it as well.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| GPT-6 Astra | Hardest math, science and reasoning | #1 LiveBench Reasoning (92.65), #1 ARC-AGI-2 (95%), #1 ARC-AGI-3 | Needs a paid ChatGPT plan; 5x GPT-6.1 Sol's price | $10 / $50 |
| Claude Opus 5 | Long agentic reasoning chains | 1673 on AA-Briefcase v1.1 and 1708 on GDPval-AA v2.1, fourth among distinct models on both; 30.2% ARC-AGI-3 | Astra now leads ARC-AGI-3; 61% hallucination rate; now a legacy model | $5 / $25 |
| GPT-5.6 Sol | Value STEM flagship | #4 ARC-AGI-2 (92.5%); #5 LiveBench Reasoning (91.65), Mathematics 96.20 | FrontierMath still unpublished; scheming flagged by METR | $100/mo ChatGPT Pro; API $4 / $20 |
| GPT-6.1 Sol | Value pick, near-Astra reasoning | #2 LiveBench Reasoning (92.63), #2 ARC-AGI-2 (94.2%), #3 LiveBench Mathematics (96.83); Intelligence Index 51.8 at $0.72 an index task | 52.7% on ARC-AGI-3 against Astra's 62.7%; in ChatGPT Work and Codex, not in Chat | $2 / $10 |
| Qwen 3.7 Max | Competition math on a budget | 97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day | API-only | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Fable 5 | Math inside a coding workflow | #3 on Arena's math subcategory (1522), behind Gemini 4 Argon and Claude Opus 5; 96.0 LiveBench Mathematics | Priciest option here | $10 / $50 |
| GLM 5.3 Flash | Open-weight problem solving | Intelligence Index 41.8 under MIT, almost eight points above GLM-5.2 at less than half the size; 320B/18B active, 1M context | Needs a multi-GPU server, not a workstation; GLM 5.3 scores higher on Z.ai's own figures and is downloadable since August 28, but at more than twice the size and under a bespoke licence | Open weights (MIT) |
GPT-6 Astra keeps the problem-solving crown, by less than a point. ARC Prize has now scored GPT-6.1 Sol, and it takes 94.2% on ARC-AGI-2, second to Astra's 95% and ahead of Claude Opus 5.5's 93.3%, for under a quarter of Astra's cost per task. Together with its 92.63 on LiveBench Reasoning, 0.02 behind Astra, that makes GPT-6.1 Sol a near co-leader, but Astra still leads both boards and ARC-AGI-3, so the crown does not move. GPT-5.6 Sol drops to fourth on ARC-AGI-2. Gemini 4 Argon, announced September 30, leads Arena's math category on a few hundred votes but has no LiveBench or ARC Prize score yet, and Claude Fable 5, which led that Arena category, is now third. In September, Astra took this crown from GPT-5.6 Sol after its September 3 launch, and LiveBench Mathematics passed to Claude Opus 5.5.
Best AI Agent
The best AI agent is Gemini Spark if you want the agent in the cloud and Claude Cowork if you want it on your desktop. These are joint picks rather than a first and a second: Spark runs in a Google Cloud VM, so it keeps working while your laptop is shut, wired into Gmail, Docs and, since early September, Google Photos; Cowork runs on your Mac or Windows machine and drives your local apps and your screen. Spark is no longer the only agent that keeps working in the cloud. SpaceXAI's Grok Bot (August 11) gives each account a persistent cloud computer that signs into your apps, from a $20/month Cursor Pro plan; Meta's Muse (September 8) runs errands from a cloud VM of its own and is free up to 100 million tokens a week; and OpenAI dots (September 29) run on GPT-6 Astra with their own cloud computer and browser, from the $100/month Pro plan, which does not include them in the EEA, UK or Switzerland. No independent board rates these products against each other, so the choice is about where the work has to happen rather than which one scores higher. Price is no longer a tiebreaker either: Spark now reaches the $19.99/month Google AI Pro tier in the US and more than 160 other countries, and Cowork is included at no extra charge on every paid Claude plan from $20/month. ChatGPT Codex Mobile (May 14) is the alternative for coding-agent work. Read the full Gemini Spark vs Claude Cowork comparison.
| Agent | Best For | Where It Runs | Strength | Price |
|---|---|---|---|---|
| Gemini Spark | 24/7 cloud tasks, Workspace workflows | Google Cloud VM (always-on) | Always-on, deep Workspace integration | From $19.99/mo Google AI Pro |
| Claude Cowork | Desktop, app-driving, design + code | Your Mac/Windows desktop | Drives local apps, sees your screen | $20/mo Claude Pro |
| ChatGPT Codex Mobile | Coding agent on phone | OpenAI cloud + iOS/Android | Approve diffs and redirect work from phone | Included in ChatGPT plans |
| Grok Bot | Browser jobs across logged-in apps | One persistent SpaceXAI cloud computer shared by all your Bots | Signs into apps with your own logins and keeps going with your laptop closed | From $20/mo Cursor Pro, or a linked SuperGrok plan |
| OpenAI dots | Ongoing responsibilities, not one-off tasks | OpenAI cloud computer and browser | Runs on GPT-6 Astra; keeps working with your computer off; reachable in ChatGPT, Slack and by voice | From $100/mo ChatGPT Pro (not EEA/UK/CH); Business Premium |
| Meta Muse | Free personal errands | A Muse Secure VM in Meta's cloud | Free up to 100M tokens a week; a separate Sentinel agent approves what it sends out | Free; $20/mo Power, $100/mo Maximum |
Gemini Spark (cloud) and Claude Cowork (desktop) remain joint picks, because this category crowns products and no new agent product shipped since the last update. The model layer moved. Gemini 4 Argon, announced September 30, leads Artificial Analysis's AutomationBench at 77.5%, ahead of Claude Sonnet 5.5 at 71.3% and Claude Opus 5.5 at 69.5%, and enters Arena's Agent Arena eighth, with a 14.2% confirmed-success rate level with Opus 5.5's 14.0%. It is not yet available outside Google's Fairwind program, though, and Claude Opus 5.5 still leads both of Artificial Analysis's agentic boards, AA-Briefcase v1.1 at 1822 and GDPval-AA v2.1 at 1846, with Sonnet 5.5 second on both. Claude Fable 5.1 still tops the Agent Arena, with 17.5% confirmed success. Among open models there, Tencent's Hy4 Preview is twelfth, DeepSeek V4.1-Flash fifteenth and Kimi K3 sixteenth. In September, Grok Bot, Meta's Muse and OpenAI dots joined Gemini Spark as always-on cloud agents, and Spark reached the $19.99/month Google AI Pro tier.
The leading models like ChatGPT, Claude and Gemini, together on Mac, iPhone and iPad.
Free to start, 4.7★ across 27,000+ reviews.
Download Fello AIBest AI for Students
The best AI for students is GPT-5.6 Luna inside ChatGPT for general coursework and Gemini 3.6 Flash inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Sonnet 5.5 as the free alternative for essay editing. Most students don't need to pay, and the free tier improved on August 6: GPT-5.6 Luna became the default for ChatGPT Free and Go, with unlimited text chats and a Think button for harder questions, though file uploads and image tools stay capped. On September 22 OpenAI added GPT-6 Luna for Free and Go users in the ChatGPT desktop app, a current-generation model on the free tier from launch day, though it scores the same 37.3 as GPT-5.6 Luna on Artificial Analysis's index. Gemini 3.6 Flash is still what the free Gemini app serves, Claude Sonnet 5.5, released September 28, is free on claude.ai, and DeepSeek V4 is free on DeepSeek's chat site. Luna is the cheapest member of the GPT-5.6 family rather than the strongest, so reach for the Think button or switch to GPT-5.5 when an essay needs more care. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's paid flagship and GPT-6 Astra now reports 97.6% on FrontierMath Tier 4 v2 against GPT-5.5 Pro's verified 39.6%, though Astra is not on ChatGPT's free tier; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).
| Task | Best Model | Why | Free? | Alternative |
|---|---|---|---|---|
| Essays & coursework | GPT-5.6 Luna | Free default in ChatGPT since August 6; unlimited text chats, plus GPT-6 Luna in the desktop app since September 22 | Yes | Claude Sonnet 5.5 (free Claude) |
| STEM problem-solving | GPT-5.6 Sol / Qwen 3.7 Max | Paid STEM flagship; Astra reports 97.6% FrontierMath Tier 4 v2 / 97.1 HMMT 2026 Feb | Pro paid / Qwen API paid | Gemini 3.6 Flash (free) |
| Research & accuracy | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, native Google Search grounding | Yes (Gemini app) | Claude Opus 5 |
| Writing editing | Claude Sonnet 5.5 | Free on claude.ai since September 28 and second only to Opus 5.5 on the Intelligence Index; Claude Fable 5.1 is the quality leader | Yes (Claude free) | Claude Fable 5.1 |
| Multimodal study (PDFs, slides, images) | Gemini 3.6 Flash | 1M context, free in Gemini app | Yes | NotebookLM (Google) |
Nothing a student can use for free changed since the last update: GPT-5.6 Luna is still the free ChatGPT default, Gemini 3.6 Flash still runs the free Gemini app, and Claude Sonnet 5.5 is still free on claude.ai. Gemini 4 Argon, announced September 30, is not available to students at all yet, and Google has said nothing about bringing it to the free Gemini app or the Google AI Pro plan. In September, Claude Sonnet 5.5 replaced Sonnet 5 as the free writing and editing pick, and OpenAI added GPT-6 Luna for Free and Go users in the ChatGPT desktop app, a cheaper model that scores the same as GPT-5.6 Luna.
Best AI for Work & Professionals
The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5.5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark is the always-on cloud agent that lives inside Google Workspace, and since July 30 it reaches the $19.99/month Google AI Pro tier as well as Google AI Ultra; OpenAI dots, Grok Bot and Meta's Muse are the alternatives, covered in the agents section above.
| Use Case | Best Model | Key Stat | Price | Alternative |
|---|---|---|---|---|
| Daily knowledge work | GPT-5.6 | ChatGPT's default since July 9; the assistant most people can open | $20/mo ChatGPT Plus | Claude Opus 5.5 |
| ChatGPT Work & Codex | GPT-6.1 Sol | In ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu since September 29, at $2 / $10 in the API | $20/mo ChatGPT Plus | GPT-6 Luna for volume work |
| Coding (proprietary) | Claude Opus 5.5 | #1 Arena WebDev (1818) and #1 Intelligence Index (57.6), at $4 / $20 | $20/mo Claude Pro | Claude Sonnet 5.5, which leads Terminal-Bench 4.0 at half the token price |
| Coding (cost-effective) | Claude Sonnet 5.5 | #1 Terminal-Bench 4.0 (63.6%) and #1 LiveBench Coding (91.4); $1.08 an index task at high effort | $2 / $10 | Qwen3.8-Max-0902; DeepSeek V4.1-Flash |
| Research & briefings | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, Google grounding | Google AI Pro / Ultra | Claude Opus 5.5 |
| Hard math, physics, finance modelling | GPT-6 Astra | #1 LiveBench Reasoning and ARC-AGI-2 (95%); reports 97.6% FrontierMath Tier 4 v2 | $100/mo ChatGPT Pro | GPT-5.6 Sol; Qwen 3.7 Max |
| Always-on agent workflows | Gemini Spark | Always-on cloud agent inside Google Workspace | From $19.99/mo Google AI Pro | Claude Cowork |
| Live news, X-context creative | Grok 4.7 | Intelligence Index 46.3 + native X grounding; the Grok app still serves 4.6 | $30/mo SuperGrok | Gemini 3.1 Pro |
| All-in-one consolidation | Fello AI | ChatGPT + Claude + Gemini + Grok + DeepSeek | $9.99/mo | Pay each vendor separately |
No row in this guide changes hands. Gemini 4 Argon, announced September 30, leads the Vals Index of finance, legal and tax work and Artificial Analysis's AutomationBench, which makes it the model to watch for professional work, but only Google's Fairwind partners can use it, with paid API customers and Google AI Ultra subscribers next and no date. In September, Claude Opus 5.5 took the proprietary coding row at a lower price than Opus 5, Claude Sonnet 5.5 took the cost-effective coding row, and GPT-6.1 Sol took the ChatGPT Work and Codex row on September 29.
AI Model Pricing in October 2026
From $0 free tiers to $199.99/month Google AI Ultra. The most consequential price on this table is Claude Opus 5.5 at $4 / $20, because on September 22 Anthropic shipped the highest-scoring model any independent evaluator has measured at a lower list price than the model it replaces: Claude Opus 5 is $5 / $25, and cache reads fall 60% to $0.20, with Anthropic's own figure putting a typical workload 40% cheaper than on Opus 5. That inverts the shape this section has had all year, in which the best model was also the most expensive. Claude Fable 5.1 and GPT-6 Astra both list at $10 / $50 and both now score below it. Six days later Anthropic added Claude Sonnet 5.5 at $2 / $10, second only to Opus 5.5 on the index, but check the effort setting before counting the saving: at max effort an index task costs $7.60, more than Opus 5.5's $5.98, while at high effort it costs $1.08. Meta's Muse Spark 1.3 lists at $1.25 / $4.25, unchanged across three capability bumps since July, with a Contributor tier at $0.10 / $0.20 for anyone willing to let Meta train on their prompts, and Grok 4.6 and Grok 4.7 both list at $2 / $6, with the caveat that a prompt of 200,000 tokens or more re-bills the entire request at $4 / $12. Grok 4.7 is the clearest case on this page of a price that did not move while the cost did: the tokens cost the same, and an Intelligence Index task costs $2.73 against Grok 4.6's $1.86, because 4.7 emits roughly twice the output getting there. The cheap end has moved twice. DeepSeek raised both V4 models roughly fourfold on August 16, which cost V4-Flash the price-performance pick, then reversed most of it on September 10: V4-Flash was retired and V4.1-Flash took its place at $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak. Among models you can buy by the token the pick stays with GLM 5.3 Flash, now at its plain $0.15 / $0.50 list since the launch promotion expired on September 9, because Artificial Analysis measures it at $0.25 an Intelligence Index task for a score of 41.8 against DeepSeek's $0.27 for 39.5. Off-peak the two are level on input and GLM is cheaper on output; at peak GLM wins both. One open model beats them both on cost per task and it was published on September 21: Xiaomi's MiMo-V2.6-Pro runs an index task for $0.13 at a score of 46.3, under plain MIT, but you have to host 1.02 trillion parameters to get it. At the frontier the value pick is Meta's Muse Spark 1.3, at $1.60 an index task for 48.1. Among closed models GPT-6 Luna took that title on September 22, at $0.10 / $0.50 by the token and $0.07 an index task, the cheapest per task of anything on this page; GPT-6 Sol fell the same way, from $4 / $20 to $2 / $10, landing exactly on Claude Sonnet 5.5's rate, and OpenAI says both cuts are permanent rather than promotional. GPT-6.1 Sol arrived on September 29 at that same $2 / $10 with cached input at $0.10, and at $0.72 an index task it is the cheapest route to a top-six Intelligence Index score. Google's Gemini 4 Argon, announced September 30, launches at an introductory $2 / $10 with cached input 95% off, which puts an index task at $1.99, though the standard rate is $4 / $20 and only Google's Fairwind partners can buy it so far. For a deeper breakdown see our full AI Pricing Comparison Guide.
| Model | Input (per 1M) | Output (per 1M) | Context | Free Access? |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 | 1,050,000 (922,000 max input) | In Chat on Pro, Business and Enterprise since September 4; Plus gets it in ChatGPT Work and Codex rather than Chat; no free tier; live in the API |
| GPT-6 Sol | $2.00 | $10.00 | 1,050,000 (922,000 max input) | ChatGPT Work & Codex on Plus, Pro, Business, Enterprise and Edu since September 22; API; not in Chat |
| GPT-6.1 Sol | $2.00 | $10.00 | 1,050,000 (922,000 max input) | ChatGPT Work & Codex on Plus, Pro, Business, Enterprise and Edu since September 29; API; not in Chat; cache reads $0.10 |
| GPT-6 Luna | $0.10 | $0.50 | 1,050,000 (922,000 max input) | Free and Go in the ChatGPT desktop app; ChatGPT Work & Codex on paid plans; API; not in Chat |
| GPT-5.5 | $5.00 | $30.00 | 1M (400K in Codex) | ChatGPT Free; API paid |
| GPT-5.5 Pro | $30.00 | $180.00 | 1M | ChatGPT Pro from $100/mo ($200 and $500 higher-usage tiers) |
| GPT-5.6 Sol | $4.00 | $20.00 | Not published | Live in ChatGPT, Codex & API (July 9); promotional rate through at least November 21, 2026 |
| GPT-5.6 Terra | $2.00 | $12.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Luna | $0.20 | $1.20 | Not published | Live in ChatGPT, Codex & API (July 9) |
| Claude Opus 5.5 | $4.00 | $20.00 | 1M | Claude Pro/Max/Team; API paid; cache reads $0.20 |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Legacy model at Anthropic; Pro/Max/API |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Legacy model at Anthropic; Pro/Max/API |
| Claude Fable 5.1 | $10.00 | $50.00 | 1M | Cache reads $0.25; in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits. Mythos 5.1 bills identically, by invitation only |
| Claude Fable 5 | $10.00 | $50.00 | 1M | Permanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits |
| Claude Sonnet 5.5 | $2.00 | $10.00 | 1M | Free on claude.ai; Pro/Max/Team; API paid; cache reads $0.20 |
| Claude Sonnet 5 | $2.00 | $10.00 | 1M | Legacy model since September 28; the fallback for flagged cyber requests; API paid |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | API paid (superseded by Sonnet 5) |
| Gemini 3.1 Pro | $2.00 (≤200K) / $4.00 (>200K) | $12.00 (≤200K) / $18.00 (>200K) | 1M | Limited Gemini app; API paid |
| Gemini 4 Argon | $2.00 intro / $4.00 standard ($0.10 cached, intro) | $10.00 intro / $20.00 standard | 1M (1M output) | Fairwind Program partners only; paid API and Google AI Ultra next, no date |
| Gemini 3.8 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | AI Studio, Antigravity, Gemini Enterprise + paid API; in the Gemini app reported for AI Pro/Ultra |
| Gemini 3.7 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | AI Studio, Android Studio, Antigravity + paid API; in the Gemini app via Spark only (AI Pro/Ultra, excludes EEA/UK/CH/Nigeria) |
| Gemini 3.6 Flash | $0.75 intro / $1.50 from Jan 1, 2027 | $3.75 intro / $7.50 from Jan 1, 2027 | 1M | Free Gemini app default; AI Studio; free API tier + paid API |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | AI Studio; free API tier + paid API |
| Qwen3.8-Max-0902 | $2.00 ($0.17 explicit cache hit, $0.25 implicit) | $6.00 | 1M (128K output) | QwenCloud API only; no consumer chat front-end, no open weights |
| Qwen 3.7 Max | $1.25 promo / $2.50 list | $3.75 promo / $7.50 list | 1M | 200 free requests/day; API paid beyond that |
| MiniMax M3 | $0.30 (50% off $0.60) | $1.20 (≤512K) | 1M | Open weights; hosting costs apply |
| LongCat-2.0 | Provider-dependent | Provider-dependent | 1M | Open weights (MIT); hosting costs apply |
| NVIDIA Nemotron 3 Ultra | Provider-dependent | Provider-dependent | 1M | Open weights (OpenMDW); hosting costs apply |
| Qwen 3.5 (open-weight) | Self-host / Together | Self-host / Together | 1M | Open weights; hosting costs apply |
| Nex-N2-Pro | Self-host / providers | Self-host / providers | 1M | Open weights (Apache 2.0); hosting costs apply |
| Rio 3.5 Open 397B | Self-host / providers | Self-host / providers | 1M | Open weights (MIT); hosting costs apply |
| Grok 4.3 | $1.25 | $2.50 | 1M | Free consumer plan; API paid |
| Grok 4.6 | $2.00 (<200K) / $4.00 (≥200K) | $6.00 (<200K) / $12.00 (≥200K) | 500K | SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel, Cloudflare |
| Grok 4.7 | $2.00 (<200K) / $4.00 (≥200K) | $6.00 (<200K) / $12.00 (≥200K) | 500K | SpaceXAI API, Cursor, Grok Build, harnesses and cloud routers; not in the Grok app |
| Muse Spark 1.3 | $1.25 ($0.10 Contributor tier) | $4.25 ($0.20 Contributor tier) | 1M | Paid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount |
| Kimi K3 | $3.00 ($0.30 cache-hit) | $15.00 | 1M | Free basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License) |
| Gemini Omni Flash (video) | $1.50 | $17.50 (video output) | 10-second clips | Gemini app / Flow; AI Studio + API |
| DeepSeek V4-Pro | $0.66 off-peak / $1.32 peak ($0.022 cache-hit off-peak) | $1.98 off-peak / $3.96 peak | 1M | DeepSeek Chat free; API paid, peak/off-peak tiers since August 16 |
| DeepSeek V4.1-Flash | $0.15 off-peak / $0.30 peak ($0.003 cache-hit off-peak) | $0.60 off-peak / $1.20 peak | 1M | DeepSeek Chat free; API paid, peak/off-peak tiers; replaced V4-Flash on September 10 |
| Kimi K2.7 Code | Provider-dependent | Provider-dependent | 256K | Open weights; hosting costs apply |
| GLM-5.2 | Provider-dependent | Provider-dependent | 1M | Open weights; hosting costs apply |
| GLM 5.3 | $1.40 ($0.26 cache-hit) | $4.40 | 1M | Z.ai API, GLM Coding Plan and ZCode; weights on Hugging Face since August 28 under the bespoke GLM 5.3 License |
| GLM 5.3 Flash | $0.15 ($0.03 cache-hit) | $0.50 | 1M | Z.ai API, GLM Coding Plan, OpenRouter; MIT weights on Hugging Face; the launch promotion ended September 9 |
| Tencent Hy4 Preview | $0.834 ($0.042 cache-hit) | $2.501 | 1M+ | Open weights (Apache 2.0); Tencent Cloud TokenHub, OpenRouter, WorkBuddy, CodeBuddy |
| Qwen3.8-Flash-Next | Provider-dependent | Provider-dependent | 262K (to 1M) | Open weights (Qwen Community License 1.0); hosting costs apply |
| MiMo-V2.6-Pro | $0.435 | $0.87 | 1M | Open weights (MIT); hosting costs apply |
| ERNIE 5.1 | China-region pricing | China-region pricing | 256K | Baidu free tier |
| Gemini Spark (agent) | Not API-priced | Not API-priced | 1M (Gemini base) | Google AI Pro $19.99, AI Ultra $99.99 or $199.99/mo |
| Fello AI (aggregator) | Routed via app | Routed via app | Model-dependent | $9.99/mo, free tier available |
The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. The GPT-6 tiers carry the same surcharge past 272K input tokens, applied to the whole request, and their cached input reads are 90% off: $1.00 for Astra, $0.20 for GPT-6 Sol and $0.01 for Luna, with GPT-6.1 Sol at 95% off, $0.10. Grok 4.6 works the same way but bites harder: at 200,000 prompt tokens and above, SpaceXAI re-bills the entire request at the higher rate rather than only the overage, so crossing the line by a thousand tokens doubles the cost of the whole call. DeepSeek is the other rate card that needs reading twice: since August 16 both V4 models bill at peak rates from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays, and at exactly half that in every other hour, so the same job can cost twice as much depending on when you run it. If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.
Best Open-Weight Models in October 2026
The best open-weight model in October 2026 is MiMo-V2.6-Pro, which Xiaomi published on September 21 under plain MIT: 1.02 trillion parameters with 42 billion active, a 1M-token context, and the weights on Hugging Face on the day. Artificial Analysis measures it at 46.3 on Intelligence Index v4.3.2, which is the highest open score it has ever published, above GLM-5.3 at 44.8 and Kimi K3 at 43.6, and it does that at $0.13 an Intelligence Index task, cheaper than any other open model on this page. It also takes 34.8% on Terminal-Bench 4.0 and 1678 on GDPval-AA v2.1, which puts it above Kimi K3 on both. Two honest limits: Arena's human voters are cool on it, 23rd on WebDev and 25th on text, and it has barely ten days of independent track record. Kimi K3 is still the highest-placed open model on Arena, twelfth on WebDev at 1658, though on the Agent Arena it now sits sixteenth, behind Tencent's Hy4 Preview at twelfth and DeepSeek V4.1-Flash at fifteenth, and GLM-5.3 is still the strongest open model on Artificial Analysis's Terminal-Bench 4.0 at 41.9% against MiMo's 34.8%, on a bespoke Z.ai licence rather than MIT. The practical recommendation for teams running their own weights is still GLM 5.3 Flash (Z.ai, MIT), released on August 26 and scoring 41.8 on v4.3.2, 320B total with 18B active, because a 1.02-trillion-parameter model is not a workstation proposition and the K3 download is 96 safetensors shards and about 1.56 TB. The cheap open tier changed again on September 10: DeepSeek retired V4-Flash 0731 and replaced it with DeepSeek-V4.1-Flash, 552B with 8B active at prefill and 16B at decode, natively multimodal, MIT on the day, scoring 39.5 on v4.3.2 against 34.3 for the build it replaces, and priced back down to $0.15 / $0.60 off-peak. It also has the highest confirmed task-success rate of any open model on the Agent Arena, 12.1%, though it ranks fifteenth there overall. Of the three releases that closed August, GLM 5.3 published its weights on August 28 under a bespoke licence with a clause requiring any model-as-a-service operator above $10 billion in revenue to pass a Z.ai security review first; Alibaba's Qwen3.8-Flash-Next, a 125B mixture-of-experts with only 6B active that previews the Qwen4 architecture, scores 39.8 and sits fourteenth on Arena's WebDev board; and Tencent's Hy4 Preview, 770B total and 49B active under Apache 2.0, has no Artificial Analysis rating but sits sixteenth on Arena WebDev at 1633. Note also that GLM 5.3 Flash is not a trimmed GLM 5.3 but a separate model on a newly trained base, which is why it could ship under plain MIT while the flagship could not.
| Model | Best For | Key Benchmark | Context / License | Where To Run |
|---|---|---|---|---|
| MiMo-V2.6-Pro | Highest open Intelligence Index ever measured | II 46.3 (v4.3.2), above GLM-5.3 and Kimi K3, at $0.13 an index task; 34.8% Terminal-Bench 4.0; GDPval-AA v2.1 1678; 1.02T/42B active | 1M / MIT | Hugging Face (XiaomiMiMo), providers, self-host |
| Kimi K3 | Highest-placed open model on Arena | II 43.6 (v4.3.2); #12 Arena WebDev, the best open placement there; #16 Agent Arena; 2.8T/104B active | 1M / Kimi K3 License | Hugging Face (96 shards, 1.56 TB), Moonshot API, providers |
| GLM 5.3 Flash | Best open model most teams can actually run | II 41.8 (v4.3.2), on the intelligence-versus-cost Pareto frontier at about $0.25 an index task; 320B/18B active, natively multimodal | 1M / MIT | Hugging Face, Z.ai API ($0.15/$0.50), OpenRouter |
| GLM-5.2 | Fallback with an independent track record | II 33.7 (v4.3.2); #25 Arena WebDev (1,604.6), #18 Agent Arena; 744B/40B active | 1M / MIT | Z.ai, Hugging Face, OpenRouter |
| GLM 5.3 | Open since August 28, but on a bespoke licence | II 44.8 (v4.3.2), the second-highest open score we found; same 744B base as GLM-5.2, post-trained only, published checkpoint counted at 753B; CyberGym 84.5% and Terminal-Bench 3.0 28.3 (vendor figures) | 1M / GLM 5.3 License (MaaS above $10B revenue needs a Z.ai security review) | Hugging Face, Z.ai API ($1.40/$4.40), GLM Coding Plan, ZCode |
| DeepSeek V4.1-Flash | Best open value if you host it yourself | II 39.5 (v4.3.2) against 34.3 for the V4-Flash 0731 build it replaces; 552B, 8B active at prefill and 16B at decode | 1M / MIT | DeepSeek API ($0.15/$0.60 off-peak, $0.30/$1.20 peak since September 10), local |
| Tencent Hy4 Preview | Largest permissively licensed model yet | 770B/49B active; 2.99/4.00 on Tencent's own 203-task expert panel vs Kimi K3 2.94 and GLM 5.3 2.92 (vendor); no Artificial Analysis rating; #16 Arena WebDev (1633), #12 Agent Arena | 1M+ / Apache 2.0 | Hugging Face (BF16 and FP8), Tencent Cloud TokenHub, OpenRouter |
| Qwen3.8-Flash-Next | Preview of the Qwen4 architecture | 125B total plus a 51B N-gram embedding, 6B active; 91.7 GPQA Diamond and 62.5 SWE-Bench Pro (vendor); II 39.8 (v4.3.2); #14 Arena WebDev (1638) | 262K to 1M / Qwen Community License 1.0 | Hugging Face, providers, self-host |
| LongCat-2.0 | Frontier open coder trained on Chinese chips | II 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active | 1M / MIT | Hugging Face, GitHub, OpenRouter |
| MiniMax M3 | Cheap frontier-class multimodal | II 44 (v4.1), 59% SWE-Bench Pro, multimodal | 1M / license TBD | Hugging Face, API $0.30/1M (50% off) |
| Nex-N2-Pro | Strongest open coding score | II 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B active | Qwen-based / Apache 2.0 | Hugging Face, providers, self-host |
| Kimi K2.7 Code | Strongest commercially-licensed open coder | +21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active | 256K / Modified MIT | Hugging Face, DeepInfra, providers |
| DeepSeek V4-Pro | Agentic real-world work | II 44 (v4.1), 1.6T/49B active | 1M / MIT | DeepSeek API ($0.66/$1.98 off-peak, $1.32/$3.96 peak since August 16), local |
| Hy3 | Newest permissive-licence entrant | II 41 (v4.1); #49 on Arena WebDev (1,508.6) | Apache 2.0 | Hugging Face, providers, self-host |
| Inkling | Thinking Machines' first open model | II 41 (v4.1), agentic 32.3, released July 15 | Open weights | Hugging Face, providers, self-host |
| Inkling Small | Same family at a quarter the size | II 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio in | Apache 2.0 | Hugging Face (BF16 and NVFP4), providers, self-host |
| NVIDIA Nemotron 3 Ultra | NVIDIA-tuned, fully permissive license | II 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active | 1M / OpenMDW | OpenRouter, Hugging Face, AWS (8x B200 self-host) |
| Qwen 3.5 (397B / 17B active) | Multimodal, fast decode | 88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v6 | 1M / open | Together, OpenRouter, local |
| Qwen3.6-35B-A3B | Efficient open agentic coder (3B active) | 86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active | 262K (to 1M YaRN) / Apache 2.0 | Hugging Face, OpenRouter, local |
| Qwen3.6-27B | Laptop-runnable dense coder | 87.8 GPQA Diamond, dense 27B, multimodal | 256K / Apache 2.0 | Local Mac/PC, Hugging Face, OpenRouter |
| Rio 3.5 Open 397B | Qwen 3.5 fine-tune, multilingual reasoning | 70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5 | 397B/17B active / MIT | Hugging Face, providers, self-host |
| Llama 4 Maverick | Meta-line flagship | 17B active / 400B total params | Llama 4 license | Meta cloud, Hugging Face, local |
| NVIDIA Nemotron 3 Nano Omni | Edge / low-power | Multimodal, very small footprint | Compact / open | Local, NVIDIA tool |
Licensing matters as much as raw score here, and August widened the gap between the two groups. Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM 5.3 Flash (MIT), MiMo-V2.6-Pro (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Hy4 Preview (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use with no revenue test. The rest carry conditions worth reading before you build on them: MiniMax M3 and MiniMax H3 ship under their own community licences, H3 additionally requiring an application from the US, EU, UK and South Korea; Kimi K3 sits on a custom licence with a revenue threshold for anyone reselling it as a service; GLM 5.3 requires a Z.ai security review of any model-as-a-service operator above $10 billion in revenue; and the Qwen Community License 1.0 on Qwen3.8-Flash-Next requires a separate licence from Qwen to run a model-as-a-service or an AI work-assistant business, though internal use is exempt. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.
Benchmarks, Prices, and Hands-On Use
Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 4.0 board, covering the Intelligence Index, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.
Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We rank on measured scores: a category crown moves as soon as a model out-scores the holder on the boards that define that category, without waiting for the vote-based boards to catch up, and a model that is not yet widely available is flagged on its card rather than denied one. We re-check ranks as well as figures, because a score can stay correct while the board around it moves. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.
Fello AI brings the top AI models together in one lightweight native app.
Switch between them anytime, no separate subscriptions.
Download Fello AI




