Best AI for Writing
Best AI for Writing: Claude Fable 5 (#1 on Arena creative writing and LiveBench Language)
The best AI for writing is Claude Fable 5, the only model in the top three of all three independent writing boards, with Claude Sonnet 5 as the free-tier value pick and GPT-5.5 as the alternative for fact-anchored business writing. Fable 5 leads Arena’s creative-writing leaderboard at 1508 Elo, tops LiveBench Language at 90.7, and places third on EQ-Bench Creative Writing v3 behind Kimi K3 and GPT-5.6 Sol. No other model is top-three on more than one of them.
This is a change from last month, when Claude Sonnet 5 held this slot on the strength of its GDPval-AA score. The preference boards do not support that placing. Sonnet 5 sits #53 on Arena creative writing and #13 on EQ-Bench, and scores 75.0 on LiveBench Language against Fable 5’s 90.7. It is still the model most people should write with day to day, because it is free and default on claude.ai, but it is the value pick rather than the quality leader.
Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits, with Pro and Team Standard reaching it through usage credits. If your writing is a work deliverable rather than prose, Claude Opus 5 tops Artificial Analysis’s GDPval-AA v2 professional-deliverables board outright at 1861, well clear of Fable 5’s 1747. Arena and EQ-Bench have not rated Opus 5 yet, so we have kept it out of the crown for now.
Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|
Claude Fable 5 | Best writing overall | #1 Arena creative writing (1508), #1 LiveBench Language (90.7), #3 EQ-Bench | Priciest option here | $10 / $50 |
Kimi K3 | Creative fiction and voice | #1 EQ-Bench Creative Writing (2377), 234 Elo clear of second | Only #10 on Arena creative writing | $3 / $15 |
Claude Sonnet 5 | Free everyday writing | Free and default on claude.ai, 1M context | #53 Arena creative writing; 75.0 LiveBench Language | $2 / $10 intro (then $3 / $15) |
Claude Opus 5 | Professional deliverables | #1 GDPval-AA v2 at 1861, ahead of Fable 5 (1747) | Not yet rated on Arena or EQ-Bench | $5 / $25 |
GPT-5.5 | Fact-anchored business writing | Documented factual-reliability gains over GPT-5.4 | Artificial Analysis now marks its reasoning tiers deprecated | $5 / $30 |
Gemini 3.6 Flash | Bulk drafts at scale | 17% fewer output tokens than 3.5 Flash | Weaker on hardest reasoning | $1.50 / $7.50 |
Runner-up and alternatives: Kimi K3 is the runner-up for creative fiction and wins EQ-Bench outright, Claude Sonnet 5 is the runner-up on value and the one to use if you are not paying, Claude Opus 5 is the pick for professional deliverables, and Gemini 3.6 Flash is the pick for bulk drafting.
What changed this month: the writing crown moved from Claude Sonnet 5 to Claude Fable 5. We widened the evidence beyond Artificial Analysis to the three boards that actually measure writing, and Fable 5 is top-three on all of them while Sonnet 5 is top-three on none. Kimi K3 (July 16) arrived as the EQ-Bench leader, and Claude Opus 5 (July 24) took GDPval-AA v2 at 1861, though the preference boards have not rated it yet.
Best AI for Chat & Daily Assistant
Best AI for Chat & Daily Assistant: GPT-5.6 (ChatGPT’s default since July 9)
The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #11 on Arena’s text leaderboard at 1485, where Claude Fable 5 leads at 1507. If you want the best-rated conversational model and are willing to leave ChatGPT, that is the swap to make.
Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 at roughly half the cost. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month for roughly 5x Plus usage or $200/month for roughly 20x), through the API (Luna $1 / $6, Terra $2.50 / $15, Sol $5 / $30 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI’s system card and the evaluator METR flagged elevated “scheming” behaviour in Sol.
GPT-5.5 is still sold by OpenAI at $5 / $30 and its Instant tier is still the safer pick for hallucination-sensitive work, with a documented 52.5% drop in hallucinated claims over GPT-5.3 Instant. Note that Artificial Analysis has since marked every GPT-5.5 reasoning tier deprecated, so treat it as a model you can still buy rather than a current benchmark reference. Claude Opus 5 is the better pick when you want a model that pushes back on weak prompts, and Gemini 3.6 Flash is the better pick if you are running everything through the free Gemini app.
Model | Best For | Strength | Weakness | Price |
|---|
GPT-5.6 | Everyday chat, ChatGPT’s default | The assistant most people can open; Terra matches GPT-5.5 at ~half cost | #11 on Arena text; scheming flagged by METR | Free / $20/mo Plus; API $1 / $6 to $5 / $30 |
Claude Fable 5 | Highest-rated conversation | #1 on Arena text overall (1507) and 6 of 7 subcategories | No free tier; usage-credit access on Pro | $10 / $50 API |
Claude Opus 5 | Thoughtful, nuanced answers | #1 Artificial Analysis Intelligence Index (61) and Agentic Index (55.3) | Not yet rated on Arena | $20/mo Pro, $5 / $25 API |
GPT-5.5 Instant | Hallucination-sensitive daily work | 52.5% fewer hallucinated claims vs 5.3 Instant | Reasoning tiers now marked deprecated by Artificial Analysis | $20/mo Plus; API $5 / $30 |
Gemini 3.6 Flash | Fast, free, multimodal | Free in the Gemini app, 1M context, #12 on Arena text | Weaker on hardest reasoning | Free / $1.50 / $7.50 API |
Fello AI | All the top models, one app | ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad | Routed via app, not direct | $9.99/mo |
Runner-up and alternatives: Claude Fable 5 is the runner-up and the actual preference leader, Claude Opus 5 is the runner-up for thoughtful daily use, Gemini 3.6 Flash is the runner-up for fast and free, and Grok 4.5 is the niche pick for live-news days. Fello AI is the natural pick if you want the top models in one Mac and iOS app for $9.99/month instead of juggling subscriptions.
What changed this month: we kept GPT-5.6 as the chat pick but changed the justification. Adding the human-preference boards showed it at #11 on Arena text rather than at the top, so the page now says plainly that this crown is about reach and default availability, not measured quality. Claude Opus 5 (July 24) replaces Opus 4.8 as the Anthropic pick at the same $5 / $25, and GPT-5.5’s “proven fallback” framing has been softened now that Artificial Analysis lists its reasoning tiers as deprecated.
Best AI for Images
Best AI for Images: ChatGPT Images 2.0 (#1 on text-to-image and image editing)
The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena’s text-to-image board at 1385 Elo and its image-editing board at 1465, and Artificial Analysis puts it first on its own image arena at 1337.7. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro.
The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta’s Muse Image is #3 on text-to-image and #2 on image editing, the strongest showing any Meta image model has managed. Google’s Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.
Model | Best For | Strength | Weakness | Price |
|---|
ChatGPT Images 2.0 | Images with readable text | #1 text-to-image (1385) and #1 image editing (1465) | Less photoreal than the Gemini image line | Included in ChatGPT Plus |
Reve 2.1 | Layout, typography, native 4K | #2 text-to-image at 1301, layout-preserving editing | Smaller ecosystem | Free / from $7.99/mo |
Muse Image | Image editing, Meta ecosystem | #3 text-to-image, #2 image editing (1402) | New, thin tooling around it | Meta AI app |
Nano Banana 2 (Gemini 3.1 Flash Image) | Photoreal portraits and products | Outranks Nano Banana Pro on both boards | Weaker on text in image | Gemini app / AI Studio |
Seedream 5.0 Pro | Multilingual text + region-precise editing | 10+ languages incl. Arabic RTL, lasso and layer editing | No independent benchmarks; copyright cloud | BytePlus / Magnific |
Midjourney v8 | Stylized art, illustration | Aesthetic baseline most artists prefer | Weaker on text in image | $10-$120/mo |
Grok Imagine | NSFW / Spicy Mode | Most permissive guardrails | Smaller model behind it | $30/mo SuperGrok |
Runner-up and alternatives: Reve 2.1 is the runner-up overall and the pick for layout and typography, Muse Image is the runner-up for editing an image you already have, and Nano Banana 2 is the photoreal pick. Grok Imagine is still the only frontier model that allows Spicy Mode adult content.
What changed this month: the crown is unchanged and independently confirmed, but the runner-ups were wrong and have been corrected. Reve 2.1 replaces Reve 2.0 as #2, Muse Image enters at #3 text-to-image and #2 image editing, and Nano Banana Pro has been demoted from “runner-up overall” to what the boards actually show, which is #8 to #11 and behind Nano Banana 2.
Best AI for Video
Best AI for Video: Gemini Omni Flash (#1 on both video leaderboards)
The best AI for video generation is Gemini Omni Flash, which leads Arena’s text-to-video board at 1527 Elo, a full 45 points clear of second place, and also tops Artificial Analysis’s video arena. It is #1 on both houses, which no other video model manages. Google reached the consumer launch on May 19, 2026 through the Gemini app, Flow and YouTube Shorts, and opened developer access on June 30 through AI Studio and the Gemini API. Read our full breakdown of Gemini Omni Flash.
Pricing runs $1.50 in and $17.50 per 1M video output tokens, which works out at roughly $0.10 per second of finished video, and it supports conversational editing so you can adjust a clip by describing the change. The one real limit is length: Omni Flash generates 10-second clips. If you need longer takes, Veo 3.1 remains the right tool inside the Gemini app, AI Studio and Vertex AI, with native audio and 1080p output.
This replaces Veo 3.1 at the top of the category. Veo 3.1 is a good model, but it is not the leading one: its best variant sits #6 on Arena’s text-to-video board, behind Omni Flash, ByteDance’s Dreamina Seedance 2.0 and Meta’s Muse Video. Google still wins this category, just with a different model than the page previously named.
Model | Best For | Strength | Weakness | Price |
|---|
Gemini Omni Flash | Best AI video overall | #1 on both video boards (1527 Arena), conversational editing | Caps at 10-second generations | ~$0.10/sec; Gemini app / AI Studio |
Dreamina Seedance 2.0 | Closest challenger | #2 text-to-video (1482) and #1 on image-to-video | ByteDance ecosystem, limited Western access | Dreamina / BytePlus |
Muse Video | Meta ecosystem video | #3 text-to-video at 1459 | Newest of the group, thin tooling | Meta AI app |
Veo 3.1 | Longer production clips | Native audio, 1080p, strong physics consistency | #6 on Arena video, not the quality leader | Google AI Pro / Ultra |
Kling 3.0 / 3.0 Turbo | Fast iteration at lower cost | Native 4K, 60fps, 15-second clips; Turbo shipped June 17 | Outside the top 16 on Arena text-to-video | From $10/mo |
Luma Ray 3 | Photoreal scenes | Strong realism for landscapes | Smaller community | Free / from $9.99/mo |
Runner-up and alternatives: Dreamina Seedance 2.0 is the runner-up overall and actually beats Omni Flash on image-to-video, Muse Video is third, and Veo 3.1 is the pick when 10 seconds is not enough. Runway is no longer listed here: the page previously named Gen-4, which has since been superseded by Gen-4.5, and we could not reproduce a top-tier placing for either across the boards we track. OpenAI retired the Sora 2 consumer app on April 26, 2026 and only the developer API remains, through September 24, 2026.
What changed this month: the video crown moved from Veo 3.1 to Gemini Omni Flash, since Omni Flash is #1 on both independent video boards while Veo 3.1 is sixth on Arena. The runner-up list was rebuilt on the same boards. Kuaishou’s current line is Kling 3.0 (February 4, 2026, native 4K at 60fps in 15-second clips) and Kling 3.0 Turbo (June 17, 2026), and no Kling model reaches the top 16 on Arena’s text-to-video board, so Kling is listed as the fast-iteration option rather than the runner-up.
Best AI for Coding
Best AI for Coding: Claude Fable 5 (#1 on Terminal-Bench 2.1, LiveBench Coding and the Remote Labor Index)
The best AI for coding is Claude Fable 5, and it holds that position on more independent boards than any other model. It is #1 on the official Terminal-Bench 2.1 leaderboard at 83.8% running in Claude Code, #1 on LiveBench Coding at 86.0, and #1 on Scale SEAL’s Remote Labor Index at 15.8, roughly 1.9x the next model on real contracted work. It also tops SEAL’s SWE Atlas Refactoring and Test Writing boards and takes #1 on Arena’s coding subcategory.
We have changed the evidence behind this crown. The widely quoted SWE-Bench Pro figure for Fable 5 comes from a vendor scaffold rather than a public board, and the public board disagrees with it: Scale SEAL’s SWE-Bench Pro leaderboards put Meta’s Muse Spark 1.1 first on both the public and private splits and do not list Fable 5 in the top six of either. The crown survives on the boards above, so that is what we cite now.
The contender to watch is Kimi K3, which takes #1 on Arena’s WebDev board at 1682, a 52-point margin over Fable 5, and #1 on Arena’s Agent board. Claude Opus 5 is the everyday-value pick at $5 / $25, half the price of Fable 5, and Anthropic’s own docs now tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for the highest-capability workloads. Read our cover of Claude Opus 5.
On price-per-result, Artificial Analysis’s Coding Index is not a Claude sweep and we will not pretend otherwise: GPT-5.6 Sol (xhigh) leads it at 78.3 with Opus 5 (max) at 78.0, a gap small enough to call a tie, with Fable 5 at 76.5 and Kimi K3 at 76.2. The cheapest serious contenders are Grok 4.5 at $2 / $6 and Gemini 3.6 Flash at $1.50 / $7.50. On open weights, GLM-5.2 (MIT) is the strongest available option and is #4 on Arena WebDev.
Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|
Claude Fable 5 | Best coding overall, long-horizon agentic | #1 Terminal-Bench 2.1 (83.8%), #1 LiveBench Coding (86.0), #1 Remote Labor Index | Priciest; Artificial Analysis Coding Index puts it 7th | $10 / $50 |
Kimi K3 | Web app building and agents | #1 Arena WebDev (1682) and #1 Arena Agent | Not yet on Terminal-Bench; weights not out | $3 / $15 |
Claude Opus 5 | Everyday-value agentic coding | Coding Index 78.0 (statistical tie for #1), Anthropic’s recommended default | Not yet rated on Arena or Terminal-Bench | $5 / $25 |
GPT-5.6 Sol | OpenAI flagship, agentic coding | #1 Artificial Analysis Coding Index at 78.3 (xhigh) | Absent from the official Terminal-Bench board; eval-gaming flagged by METR | $5 / $30 |
Muse Spark 1.1 | Cheap agentic tool use | #1 on SEAL SWE-Bench Pro (public and private) and MCP Atlas (88.1) | US-only preview | $1.25 / $4.25 |
Grok 4.5 | Cheap value coder | #4 on the official Terminal-Bench 2.1 board at 79.3% via Cursor CLI | Higher hallucination rate; EU API console still closed | $2 / $6 |
Gemini 3.6 Flash | Agent coding at scale | Intelligence Index 50, 17% fewer output tokens than 3.5 Flash | Weaker on hardest reasoning | $1.50 / $7.50 |
GLM-5.2 | Best open-weight coder | Highest open Intelligence Index (51), #4 Arena WebDev | Self-host or provider only | Open weights (MIT) |
Runner-up and alternatives: Kimi K3 is the runner-up on Arena’s web and agent boards, Claude Opus 5 is the runner-up on value at half the price, GPT-5.6 Sol is the runner-up on Artificial Analysis’s composite, and GLM-5.2 is the open-weight pick. Inside IDEs, Cursor with Claude is still the most popular pairing and Claude Code is the natural pick if you live in the terminal.
What changed this month: the crown stayed with Claude Fable 5 but the evidence behind it was replaced. We dropped the vendor-scaffold SWE-Bench Pro claim, because Scale SEAL’s public SWE-Bench Pro boards contradict it, and rebuilt the case on Terminal-Bench 2.1, LiveBench Coding, the Remote Labor Index and SWE Atlas, where Fable 5 is first. Kimi K3 (July 16) enters as the Arena WebDev and Agent leader, and Claude Opus 5 (July 24) replaces Opus 4.8 as the everyday-value Claude at the same $5 / $25. We have also stopped presenting vendor Terminal-Bench figures as leaderboard results: the official board has Grok 4.5 at 79.3%, not the 83.3% xAI reported, and does not list GPT-5.6 Sol at all.
Best AI for Creativity
Best AI for Creativity: Grok 4.5 (fewest content restrictions, native real-time X)
The best AI for unfiltered, on-trend creative work is Grok 4.5, and we want to be exact about why. This pick is about the product, not the prose quality. Grok 4.5 carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month.
It is not the best writer, and the boards are blunt about it. Grok 4.5 sits #33 on EQ-Bench Creative Writing and #41 on Arena’s creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. If you are picking on output quality alone, Claude Fable 5 wins outright at #1 on Arena creative writing, and Kimi K3 wins EQ-Bench. Choose Grok 4.5 for what it will let you make, not for how well it writes.
Model | Best For | Strength | Weakness | Price |
|---|
Grok 4.5 | Unfiltered, opinionated, on-trend | Fewest content restrictions, native real-time X grounding | #33 EQ-Bench, #41 Arena creative writing | $30/mo SuperGrok |
Claude Fable 5 | Highest-quality creative prose | #1 Arena creative writing (1508), #1 LiveBench Language | Cautious guardrails on edgy briefs | $10 / $50 API |
Kimi K3 | Fiction and distinctive voice | #1 EQ-Bench Creative Writing at 2377 | Only #10 on Arena creative writing | $3 / $15 |
Claude Opus 5 | Long-form structured creativity | Holds long threads and self-edits; #1 Intelligence Index | Most cautious of the group | $20/mo Pro, $5 / $25 API |
Gemini 3.1 Pro | Multimodal creative | Strong text, image and video chain | Quotas inside the Gemini app | Free / $2.00-$4.00 API in |
Grok Imagine (Spicy Mode) | NSFW / adult creative | Most permissive image generation | Niche use case | $30/mo SuperGrok |
Runner-up and alternatives: Claude Fable 5 is the runner-up and the right pick if quality matters more than freedom, Kimi K3 is the pick for fiction, and Claude Opus 5 is the pick for creative projects that run across many turns. For adult creative work, Grok Imagine Spicy Mode is still the only frontier-grade option.
What changed this month: Grok 4.5 keeps this pick, but the reasoning is now stated honestly. Adding the human-preference boards showed it at #33 on EQ-Bench and #41 on Arena creative writing, so this section no longer implies that it wins on quality. It is here for its permissiveness and its live X access, and the page now names Claude Fable 5 as the model to use when you want the better writing.
Best AI for Accuracy
Best AI for Accuracy: Gemini 3.1 Pro (98% on ARC-AGI-1, at $0.52 per task)
The best AI for accuracy and research is Gemini 3.1 Pro. Its strongest result is on ARC Prize’s ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you actually want when the answer has to be current rather than merely plausible.
It also scores 94.3% on GPQA Diamond and 44.4% on Humanity’s Last Exam, and tops Scale SEAL’s HLE board at 46.44. We have dropped the page’s previous ARC-AGI-2 framing. Its 77.1% is still correct, but the board has moved and that score now places it around 14th, behind GPT-5.6 Sol at 93% and Claude Opus 5 at 90%, so it is no longer evidence of an accuracy lead.
Two honest caveats. On grounded search specifically, Arena’s search leaderboard is led by Anthropic, not Google, with Gemini 3.1 Pro grounding at #7. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model according to ARC Prize. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience, which is weak ground for a crown named accuracy.
Model | Best For | Key Benchmark | Weakness | Price |
|---|
Gemini 3.1 Pro | Cheap, reliable factual work | 98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.3% GPQA | ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search | $2.00-$4.00 / $12.00-$18.00 (tiered) |
Claude Fable 5 | Grounded search | #1 and #3 on Arena’s search leaderboard, ahead of Google | No single cheap tier | $10 / $50 (Fable 5) |
GPT-5.6 Sol | Novel reasoning | #1 ARC-AGI-2 at 93%, against a 100% human panel | Scheming flagged by METR | $5 / $30 |
Claude Opus 5 | Hardest unseen problems | #1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize) | Artificial Analysis measures a 50% hallucination rate | $5 / $25 |
Qwen 3.7 Max | Frontier accuracy at value pricing | 92.4 GPQA Diamond, 200 free requests/day | API-only, no chat front-end | $1.25 / $3.75 promo; $2.50 / $7.50 list |
Claude Opus 4.6 | Honesty under pressure | #1 on Scale SEAL’s MASK board at 96.28; Anthropic holds the top 5 | Superseded as a flagship | Legacy Anthropic model |
Runner-up and alternatives: Anthropic’s models are the runner-up for grounded search and sweep the honesty-under-pressure board, GPT-5.6 Sol is the runner-up for novel reasoning, and Qwen 3.7 Max is the value pick at the frontier.
What changed this month: Gemini 3.1 Pro keeps the accuracy crown, but on rebased evidence. The lead argument is now 98% on ARC-AGI-1 at $0.52 per task plus Search grounding, and the old ARC-AGI-2 framing is gone because 77.1% now ranks around 14th rather than at the top. We also split out what this category was quietly doing at once, so grounded search now credits Anthropic and novel reasoning credits GPT-5.6 Sol and Claude Opus 5, rather than implying one model wins all three.
Best AI for Problem Solving
Best AI for Problem Solving: GPT-5.6 Sol (#1 on LiveBench Mathematics, Reasoning and ARC-AGI-2)
The best AI for hard problem solving is GPT-5.6 Sol, and it is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 93%, the closest any model has come to the 100% human panel. Three separate houses put it first on the reasoning tasks that matter, which is more agreement than any other category on this page produces.
OpenAI has still not published Sol’s FrontierMath score, so the verified OpenAI mark remains GPT-5.5 Pro’s 39.6% on FrontierMath Tier 4, and we will slot Sol’s number in the moment it goes public. Qwen 3.7 Max is the value alternative for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, at a fraction of the cost of ChatGPT Pro, and it now includes 200 free model requests per day.
Claude Opus 5 is the alternative for long agentic reasoning chains, leading Artificial Analysis’s Agentic Index at 55.3 and ARC Prize’s ARC-AGI-3 at 30%, roughly 3.75x the next-best model. It runs second to Sol on LiveBench Reasoning at 91.2. For multimodal reasoning where the problem includes diagrams or documents, Gemini 3.1 Pro is still the practical pick.
Model | Best For | Key Benchmark | Weakness | Price |
|---|
GPT-5.6 Sol | Hardest math, science and reasoning | #1 LiveBench Mathematics (96.2), #1 LiveBench Reasoning (91.7), #1 ARC-AGI-2 (93%) | FrontierMath still unpublished; scheming flagged by METR | $100/mo ChatGPT Pro; API $5 / $30 |
Claude Opus 5 | Long agentic reasoning chains | #1 Agentic Index (55.3), #1 ARC-AGI-3 (30%) | Second on LiveBench Reasoning; 50% hallucination rate | $5 / $25 |
GPT-5.5 Pro | Verified FrontierMath leader | 39.6% FrontierMath Tier 4 | Superseded by Sol as flagship | $100/mo ChatGPT Pro |
Qwen 3.7 Max | Competition math on a budget | 97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day | API-only | $1.25 / $3.75 promo; $2.50 / $7.50 list |
Claude Fable 5 | Math inside a coding workflow | #1 on Arena’s math subcategory (1543), 96.0 LiveBench Mathematics | Priciest option here | $10 / $50 |
GLM-5.2 | Open-weight problem solving | Highest open Intelligence Index at 51, MIT, 1M context | Self-host or provider only | Open weights (MIT) |
Runner-up and alternatives: Claude Opus 5 is the runner-up and the natural pick for long-chain agentic reasoning, Claude Fable 5 is the runner-up on Arena’s math board, Qwen 3.7 Max is the value pick, and GLM-5.2 is the open-weight pick.
What changed this month: GPT-5.6 Sol keeps this crown and the case for it got stronger, not weaker. Widening the research to LiveBench and ARC Prize showed it first on mathematics, reasoning and ARC-AGI-2, so this is now the most defensible pick on the page. Claude Opus 5 (July 24) enters as the agentic-reasoning alternative on the strength of ARC-AGI-3, and Qwen 3.7 Max picked up a free tier of 200 requests per day.
Best AI Agent
Best AI Agent: Gemini Spark vs Claude Cowork ($99.99/month Ultra vs $20/month Pro)
The best AI agent right now is Gemini Spark for 24/7 cloud-resident work and Claude Cowork for desktop-resident work, with ChatGPT Codex as the alternative for coding agents and OpenAI Operator-class browser agents as the alternative for web tasks. AI agents are the fastest-moving category of 2026: each top vendor now ships an agent product, and the practical choice is between agents that live in the cloud (run while your laptop is closed) and agents that live on your desktop (drive your apps directly).
Gemini Spark launched at Google I/O on May 19, 2026 and is the first 24/7 cloud agent. Claude Cowork launched in general availability on April 9, 2026 and runs as a desktop agent that drives your local apps. ChatGPT Codex Mobile (May 14) is the pick for coding-agent work, now usable from iOS and Android. Read the full Gemini Spark vs Claude Cowork comparison.
Agent | Best For | Where It Runs | Strength | Price |
|---|
Gemini Spark | 24/7 cloud tasks, Workspace workflows | Google Cloud VM (always-on) | First true 24/7 agent, deep Workspace integration | $99.99/mo Google AI Ultra |
Claude Cowork | Desktop, app-driving, design + code | Your Mac/Windows desktop | Drives local apps, sees your screen | $20/mo Claude Pro |
ChatGPT Codex Mobile | Coding agent on phone | OpenAI cloud + iOS/Android | Approve diffs and redirect work from phone | Included in ChatGPT plans |
Grok Agentic (Grok 4.5) | Real-time research, X scraping | xAI cloud | Native X integration | $30/mo SuperGrok |
OpenAI Operator-class | Browser tasks, web forms | OpenAI cloud + your browser | Web automation | ChatGPT Pro |
Runner-up and alternatives: Claude Cowork is the runner-up overall and the natural pick when you want the agent on your machine driving your apps. ChatGPT Codex Mobile is the runner-up for coding agents. Grok Agentic is the niche pick for real-time research.
What changed this month: no new consumer agents shipped, so the Gemini Spark (cloud) versus Claude Cowork (desktop) choice still drives most agent decisions for individual users. The model layer underneath them moved a lot. Claude Opus 5 (July 24) took #1 on Artificial Analysis’s Agentic Index at 55.3, ahead of GPT-5.6 Sol at 54.0 and Claude Fable 5 at 52.8, and it costs $5 / $25. Kimi K3 took #1 on Arena’s Agent board, where the metric is task success rate rather than Elo, ahead of Fable 5 and Opus 4.8. For teams building their own agents, Meta’s Muse Spark 1.1 is a cheap agent-native option at $1.25 / $4.25 that leads Scale SEAL’s MCP Atlas tool-use board at 88.1, and GLM-5.2 (MIT) is the strongest open-weight agent model at #4 on Arena’s Agent board.