Best AI for Writing
Best AI for Writing: Claude Fable 5 (#1 on Arena creative writing and LiveBench Language)
The best AI for writing is Claude Fable 5, the only model in the top three of all three independent writing boards, with Claude Sonnet 5 as the free-tier value pick and GPT-5.5 as the alternative for fact-anchored business writing. Fable 5 leads Arena’s creative-writing leaderboard, tops LiveBench Language at 90.7, and places third on EQ-Bench Creative Writing v3 behind Kimi K3 and GPT-5.6 Sol. No other model is top-three on more than one of them.
This is a change from last month, when Claude Sonnet 5 held this slot on the strength of its GDPval-AA score. The preference boards do not support that placing. Sonnet 5 sits #53 on Arena creative writing and #13 on EQ-Bench, and scores 75.0 on LiveBench Language against Fable 5’s 90.7. It is still the model most people should write with day to day, because it is free and default on claude.ai, but it is the value pick rather than the quality leader.
Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits, with Pro and Team Standard reaching it through usage credits. If your writing is a work deliverable rather than prose, Claude Opus 5 tops Artificial Analysis’s GDPval-AA v2 professional-deliverables board outright at 1858, well clear of Fable 5’s 1746. Fable 5 keeps the writing crown because it leads the boards that judge prose and factual care: Arena’s text leaderboard at 1508.6, Humanity’s Last Exam at 53.3% and AA-Omniscience at 40.
Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|
Claude Fable 5 | Best writing overall | #1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #3 EQ-Bench | Priciest option here | $10 / $50 |
Kimi K3 | Creative fiction and voice | #1 EQ-Bench Creative Writing (2377), 234 Elo clear of second | Only #10 on Arena creative writing | $3 / $15 |
Claude Sonnet 5 | Free everyday writing | Free and default on claude.ai, 1M context | #53 Arena creative writing; 75.0 LiveBench Language | $2 / $10 intro (then $3 / $15) |
Claude Opus 5 | Professional deliverables | #1 GDPval-AA v2 at 1858, ahead of Fable 5 (1746) | Behind Fable 5 on Arena text and on Humanity’s Last Exam | $5 / $25 |
GPT-5.5 | Fact-anchored business writing | Documented factual-reliability gains over GPT-5.4 | Artificial Analysis now marks its reasoning tiers deprecated | $5 / $30 |
Gemini 3.6 Flash | Bulk drafts at scale | 17% fewer output tokens than 3.5 Flash | Weaker on hardest reasoning | $1.50 / $7.50 |
Runner-up and alternatives: Kimi K3 is the runner-up for creative fiction and wins EQ-Bench outright, Claude Sonnet 5 is the runner-up on value and the one to use if you are not paying, Claude Opus 5 is the pick for professional deliverables, and Gemini 3.6 Flash is the pick for bulk drafting.
What changed this month: the writing crown stays with Claude Fable 5, and it is the strongest-supported call on the page. Fable 5 is now #1 on Arena’s text leaderboard at 1508.6 on the August 1 cutoff, #1 on Humanity’s Last Exam at 53.3% and #1 on AA-Omniscience at 40, which is three different houses agreeing. Claude Opus 5 keeps the professional-deliverables lane on GDPval-AA v2, where the re-fitted board now reads 1858 against Fable 5’s 1746.
Best AI for Chat & Daily Assistant
Best AI for Chat & Daily Assistant: GPT-5.6 (ChatGPT’s default since July 9)
The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena’s text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. If you want the best-rated conversational model and are willing to leave ChatGPT, that is the swap to make.
Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month for roughly 5x Plus usage or $200/month for roughly 20x), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $5 / $30 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI’s system card and the evaluator METR flagged elevated “scheming” behaviour in Sol.
GPT-5.5 is still sold by OpenAI at $5 / $30 and its Instant tier is still the safer pick for hallucination-sensitive work, with a documented 52.5% drop in hallucinated claims over GPT-5.3 Instant. Note that Artificial Analysis has since marked every GPT-5.5 reasoning tier deprecated, so treat it as a model you can still buy rather than a current benchmark reference. Claude Opus 5 is the better pick when you want a model that pushes back on weak prompts, and Gemini 3.6 Flash is the better pick if you are running everything through the free Gemini app.
Model | Best For | Strength | Weakness | Price |
|---|
GPT-5.6 | Everyday chat, ChatGPT’s default | The assistant most people can open; Terra matches GPT-5.5 at ~half cost | #14 on Arena text; scheming flagged by METR | Free / $20/mo Plus; API $0.20 / $1.20 to $5 / $30 |
Claude Fable 5 | Highest-rated conversation | #1 on Arena text overall (1508.6) and 6 of 7 subcategories | No free tier; usage-credit access on Pro | $10 / $50 API |
Claude Opus 5 | Thoughtful, nuanced answers | #1 Artificial Analysis Intelligence Index (61) and Agentic Index (55.3) | #6 on Arena text, behind Claude Fable 5 | $20/mo Pro, $5 / $25 API |
GPT-5.5 Instant | Hallucination-sensitive daily work | 52.5% fewer hallucinated claims vs 5.3 Instant | Reasoning tiers now marked deprecated by Artificial Analysis | $20/mo Plus; API $5 / $30 |
Gemini 3.6 Flash | Fast, free, multimodal | Free in the Gemini app, 1M context, #12 on Arena text | Weaker on hardest reasoning | Free / $1.50 / $7.50 API |
Fello AI | All the top models, one app | ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad | Routed via app, not direct | $9.99/mo |
Runner-up and alternatives: Claude Fable 5 is the runner-up and the actual preference leader, Claude Opus 5 is the runner-up for thoughtful daily use, Gemini 3.6 Flash is the runner-up for fast and free, and Grok 4.5 is the niche pick for live-news days. Fello AI is the natural pick if you want the top models in one Mac and iOS app for $9.99/month instead of juggling subscriptions.
What changed this month: GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open.
Best AI for Images
Best AI for Images: ChatGPT Images 2.0 (#1 on text-to-image and image editing)
The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena’s text-to-image board at 1385 Elo and its image-editing board at 1463, and Artificial Analysis puts it first on its own image arena at 1339.4. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro.
The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta’s Muse Image is #3 on Arena’s text-to-image board and #2 on image editing, the strongest showing any Meta image model has managed. Note that the houses disagree about third place: on Artificial Analysis’s image board it is Microsoft’s MAI-Image-2.5 at 1269.7 in third, and that model is #3 on Arena’s image-editing board too. Google’s Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.
Model | Best For | Strength | Weakness | Price |
|---|
ChatGPT Images 2.0 | Images with readable text | #1 text-to-image (1385) and #1 image editing (1463) | Less photoreal than the Gemini image line | Included in ChatGPT Plus |
Reve 2.1 | Layout, typography, native 4K | #2 text-to-image at 1302, layout-preserving editing | Smaller ecosystem | Free / from $7.99/mo |
Muse Image | Image editing, Meta ecosystem | #3 text-to-image, #2 image editing (1407) | New, thin tooling around it | Meta AI app |
Nano Banana 2 (Gemini 3.1 Flash Image) | Photoreal portraits and products | Outranks Nano Banana Pro on both boards | Weaker on text in image | Gemini app / AI Studio |
Seedream 5.0 Pro | Multilingual text + region-precise editing | 10+ languages incl. Arabic RTL, lasso and layer editing | No independent benchmarks; copyright cloud | BytePlus / Magnific |
Midjourney v8 | Stylized art, illustration | Aesthetic baseline most artists prefer | Weaker on text in image | $10-$120/mo |
Grok Imagine | NSFW / Spicy Mode | Most permissive guardrails | Smaller model behind it | $30/mo SuperGrok |
Runner-up and alternatives: Reve 2.1 is the runner-up overall and the pick for layout and typography, Muse Image is the runner-up for editing an image you already have, and Nano Banana 2 is the photoreal pick. Grok Imagine is still the only frontier model that allows Spicy Mode adult content.
What changed this month: the image crown is unchanged and GPT Image 2 still leads all three boards we track, on Arena text-to-image (1385), Arena image editing (1463) and Artificial Analysis’s image arena (1339.4). The one addition is Microsoft’s MAI-Image-2.5, which sits third on Artificial Analysis’s image board at 1269.7 and third on Arena’s image-editing board, so third place now depends on which house you read.
Best AI for Video
Best AI for Video: Gemini Omni Flash (#1 on both video leaderboards)
The best AI for video generation is Gemini Omni Flash, which leads Arena’s text-to-video board at 1527 Elo, a full 45 points clear of second place, and also tops Artificial Analysis’s video arena. It is #1 on both houses, which no other video model manages. Google reached the consumer launch on May 19, 2026 through the Gemini app, Flow and YouTube Shorts, and opened developer access on June 30 through AI Studio and the Gemini API. Read our full breakdown of Gemini Omni Flash.
Pricing runs $1.50 in and $17.50 per 1M video output tokens, which works out at roughly $0.10 per second of finished video, and it supports conversational editing so you can adjust a clip by describing the change. The one real limit is length: Omni Flash generates 10-second clips. If you need longer takes, Veo 3.1 remains the right tool inside the Gemini app, AI Studio and Vertex AI, with native audio and 1080p output.
This replaces Veo 3.1 at the top of the category. Veo 3.1 is a good model, but it is not the leading one: its best variant sits #6 on Arena’s text-to-video board, behind Omni Flash, ByteDance’s Dreamina Seedance 2.0 and Meta’s Muse Video. Google still wins this category, just with a different model than the page previously named.
The contender to watch is MiniMax H3, launched July 31, 2026. It answers Omni Flash’s two weakest points directly, generating 2K clips of 4 to 15 seconds with native stereo audio against Omni Flash’s 10-second cap, and it is billed per second from $0.13 at 2K. Artificial Analysis already places it #2 on its video board at 1241.5, only 3.3 Elo behind Omni Flash. We are not moving the crown on that: the score is one day old, it comes from the single board that has not reproduced between parses, and Arena’s text-to-video vote cutoff predates H3, so it has not yet been judged by anyone’s votes. MiniMax has promised the weights in early August.
Model | Best For | Strength | Weakness | Price |
|---|
Gemini Omni Flash | Best AI video overall | #1 on both video boards (1527 Arena), conversational editing | Caps at 10-second generations | ~$0.10/sec; Gemini app / AI Studio |
MiniMax H3 | 2K clips with native audio | 4-15s at 2K, native stereo audio; #2 on AA video (1241.5) | One day old, no Arena votes yet; weights not published | $0.13/sec at 2K |
Dreamina Seedance 2.0 | Closest challenger | #2 text-to-video (1482) and #1 on image-to-video | ByteDance ecosystem, limited Western access | Dreamina / BytePlus |
Muse Video | Meta ecosystem video | #3 text-to-video at 1459 | Newest of the group, thin tooling | Meta AI app |
Veo 3.1 | Longer production clips | Native audio, 1080p, strong physics consistency | #6 on Arena video, not the quality leader | Google AI Pro / Ultra |
Kling 3.0 / 3.0 Turbo | Fast iteration at lower cost | Native 4K, 60fps, 15-second clips; Turbo shipped June 17 | Outside the top 16 on Arena text-to-video | From $10/mo |
Luma Ray 3 | Photoreal scenes | Strong realism for landscapes | Smaller community | Free / from $9.99/mo |
Runner-up and alternatives: Dreamina Seedance 2.0 is the runner-up overall and actually beats Omni Flash on image-to-video, Muse Video is third, and Veo 3.1 is the pick when 10 seconds is not enough. Runway is no longer listed here: the page previously named Gen-4, which has since been superseded by Gen-4.5, and we could not reproduce a top-tier placing for either across the boards we track. OpenAI retired the Sora 2 consumer app on April 26, 2026 and only the developer API remains, through September 24, 2026.
What changed this month: Gemini Omni Flash keeps the video crown, and it is still #1 on both boards, at 1527.5 on Arena and 1244.8 on Artificial Analysis. MiniMax H3 (July 31) enters as the contender at #2 on Artificial Analysis’s video board (1241.5), 3.3 Elo behind, and it beats Omni Flash on specification with 2K output, clips up to 15 seconds and native stereo audio. We are holding the crown because that margin is one day old, single-board, and carries no votes on Arena, whose text-to-video cutoff predates the launch.
Best AI for Coding
Best AI for Coding: Claude Opus 5 (#1 on Arena’s WebDev and image-to-WebDev boards)
The best AI for coding is Claude Opus 5, and it wins on the two boards where developers vote on the finished result rather than a script measuring a harness. It is #1 on Arena’s WebDev board at 1,702.9 and #1 on image-to-WebDev at 1,668.6, on vote cutoffs of August 1 and July 31. Those are human comparisons with published cutoffs and no scaffold variable, which is exactly what a coding crown should rest on.
Price is the second half of the argument. Opus 5 runs $5 / $25 per 1M tokens against Claude Fable 5’s $10 / $50, and Artificial Analysis measures it at $2.34 per index task against Fable 5’s $3.15. Anthropic’s own docs tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for workloads that need the highest available capability, which is the same split. Claude Fable 5 is the runner-up and stays the pick for the hardest long-horizon work, at #4 on WebDev (1,630.7) and #2 on image-to-WebDev (1,625.7).
The contender is Kimi K3, which takes #2 on WebDev at 1,675.5, 45 points clear of Fable 5 and behind only Opus 5, and #3 on Arena’s Agent board. It is the strongest open-weight coder on the boards, and at $3 / $15 it undercuts both Claude tiers, though self-hosting it means 1.56 TB of weights across 96 shards. Read our cover of Claude Opus 5.
On price-per-result, Artificial Analysis’s Coding Index is not a Claude sweep and we will not pretend otherwise: GPT-5.6 Sol (xhigh) leads it at 78.3 with Opus 5 (max) at 78.0, a gap small enough to call a tie, with Fable 5 at 76.5 and Kimi K3 at 76.2. The cheapest serious contender is now DeepSeek V4-Flash 0731 at $0.14 / $0.28, which scores 78.7% on Artificial Analysis’s Terminal-Bench v2.1 mirror, ahead of Gemini 3.6 Flash at 77.5% and behind Grok 4.5 at 81.6%. On open weights, Kimi K3 is now the strongest option on score, since its weights shipped on July 27, while GLM-5.2 (MIT) stays the one most teams can realistically host.
Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|
Claude Opus 5 | Best coding overall | #1 Arena WebDev (1,702.9) and #1 image-to-WebDev (1,668.6); $2.34 per index task | Artificial Analysis Coding Index has GPT-5.6 Sol a shade ahead | $5 / $25 |
Claude Fable 5 | Hardest long-horizon agentic work | #2 Arena image-to-WebDev (1,625.7), #4 WebDev; 1M context | Priciest; Artificial Analysis Coding Index puts it 7th | $10 / $50 |
Kimi K3 | Web app building and agents | #2 Arena WebDev (1,675.5), #3 Arena Agent, highest open Intelligence Index (57) | 1.56 TB to self-host | $3 / $15 |
GPT-5.6 Sol | OpenAI flagship, agentic coding | #1 Artificial Analysis Coding Index at 78.3 (xhigh) | Absent from the official Terminal-Bench board; eval-gaming flagged by METR | $5 / $30 |
Muse Spark 1.1 | Cheap agentic tool use | #1 on SEAL SWE-Bench Pro (public and private) and MCP Atlas (88.1) | US-only preview | $1.25 / $4.25 |
Grok 4.5 | Cheap value coder | #4 on the official Terminal-Bench 2.1 board at 79.3% via Cursor CLI | Higher hallucination rate; EU API console still closed | $2 / $6 |
Gemini 3.6 Flash | Agent coding at scale | Intelligence Index 50, 17% fewer output tokens than 3.5 Flash | Weaker on hardest reasoning | $1.50 / $7.50 |
GLM-5.2 | Best open-weight coder you can host | Intelligence Index 51, #6 Arena WebDev (1,587.1), MIT licence | Kimi K3 outscores it; self-host or provider only | Open weights (MIT) |
Runner-up and alternatives: Kimi K3 is the runner-up on Arena’s WebDev board and the pick if you want the strongest open weights, Claude Fable 5 is the runner-up for the hardest long-horizon work, GPT-5.6 Sol is the runner-up on Artificial Analysis’s composite, and GLM-5.2 is the open-weight pick for teams hosting it themselves. Inside IDEs, Cursor with Claude is still the most popular pairing and Claude Code is the natural pick if you live in the terminal.
What changed this month: the coding crown moved to Claude Opus 5. Arena finished rating it, and it came first on both coding boards, WebDev at 1,702.9 and image-to-WebDev at 1,668.6, with Claude Fable 5 fourth and second. Those are vote-based boards with published cutoffs, so the crown now rests on human comparisons rather than on any one harness, and Opus 5 does it at half Fable 5’s price. Fable 5 becomes the runner-up for the hardest long-horizon work, and Kimi K3 stays the contender at #2 on WebDev.
Best AI for Creativity
Best AI for Creativity: Grok 4.5 (fewest content restrictions, native real-time X)
The best AI for unfiltered, on-trend creative work is Grok 4.5, and we want to be exact about why. This pick is about the product, not the prose quality. Grok 4.5 carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month.
It is not the best writer, and the boards are blunt about it. Grok 4.5 sits #33 on EQ-Bench Creative Writing and #41 on Arena’s creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. If you are picking on output quality alone, Claude Fable 5 wins outright at #1 on Arena creative writing, and Kimi K3 wins EQ-Bench. Choose Grok 4.5 for what it will let you make, not for how well it writes.
Model | Best For | Strength | Weakness | Price |
|---|
Grok 4.5 | Unfiltered, opinionated, on-trend | Fewest content restrictions, native real-time X grounding | #33 EQ-Bench, #41 Arena creative writing | $30/mo SuperGrok |
Claude Fable 5 | Highest-quality creative prose | #1 Arena creative writing, #1 LiveBench Language | Cautious guardrails on edgy briefs | $10 / $50 API |
Kimi K3 | Fiction and distinctive voice | #1 EQ-Bench Creative Writing at 2377 | Only #10 on Arena creative writing | $3 / $15 |
Claude Opus 5 | Long-form structured creativity | Holds long threads and self-edits; #1 Intelligence Index | Most cautious of the group | $20/mo Pro, $5 / $25 API |
Gemini 3.1 Pro | Multimodal creative | Strong text, image and video chain | Quotas inside the Gemini app | Free / $2.00-$4.00 API in |
Grok Imagine (Spicy Mode) | NSFW / adult creative | Most permissive image generation | Niche use case | $30/mo SuperGrok |
Runner-up and alternatives: Claude Fable 5 is the runner-up and the right pick if quality matters more than freedom, Kimi K3 is the pick for fiction, and Claude Opus 5 is the pick for creative projects that run across many turns. For adult creative work, Grok Imagine Spicy Mode is still the only frontier-grade option.
What changed this month: Grok 4.5 keeps this pick, and nothing about the product moved. It is here for its permissiveness and its live X access, not for board position, and Claude Fable 5 remains the model to use when you want the better writing. One naming note for August: xAI now trades as SpaceXAI on the leaderboards, and the model is unchanged.
Best AI for Accuracy
Best AI for Accuracy: Gemini 3.1 Pro (98% on ARC-AGI-1, at $0.52 per task)
The best AI for accuracy and research is Gemini 3.1 Pro. Its strongest result is on ARC Prize’s ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you actually want when the answer has to be current rather than merely plausible.
It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity’s Last Exam, and tops Scale SEAL’s HLE board at 46.44. We have dropped the page’s previous ARC-AGI-2 framing. Its 77.1% is still correct, but the board has moved and that score now places it around 14th, behind GPT-5.6 Sol at 93% and Claude Opus 5 at 90%, so it is no longer evidence of an accuracy lead.
Two honest caveats. On grounded search specifically, Arena’s search leaderboard is led by Anthropic, not Google, with Gemini 3.1 Pro grounding at #7. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model according to ARC Prize. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience, which is weak ground for a crown named accuracy.
Model | Best For | Key Benchmark | Weakness | Price |
|---|
Gemini 3.1 Pro | Cheap, reliable factual work | 98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol) | ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search | $2.00-$4.00 / $12.00-$18.00 (tiered) |
Claude Fable 5 | Grounded search | #1 and #3 on Arena’s search leaderboard, ahead of Google | No single cheap tier | $10 / $50 (Fable 5) |
GPT-5.6 Sol | Novel reasoning | #1 ARC-AGI-2 at 93%, against a 100% human panel | Scheming flagged by METR | $5 / $30 |
Claude Opus 5 | Hardest unseen problems | #1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize) | Artificial Analysis measures a 50% hallucination rate | $5 / $25 |
Qwen 3.7 Max | Frontier accuracy at value pricing | 92.4 GPQA Diamond, 200 free requests/day | API-only, no chat front-end | $1.25 / $3.75 promo; $2.50 / $7.50 list |
Claude Opus 4.6 | Honesty under pressure | #1 on Scale SEAL’s MASK board at 96.28; Anthropic holds the top 5 | Superseded as a flagship | Legacy Anthropic model |
Runner-up and alternatives: Anthropic’s models are the runner-up for grounded search and sweep the honesty-under-pressure board, GPT-5.6 Sol is the runner-up for novel reasoning, and Qwen 3.7 Max is the value pick at the frontier.
What changed this month: Gemini 3.1 Pro keeps the accuracy crown, but its headline number is no longer a lead. Its GPQA Diamond score re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly, so the crown rests on the ARC-AGI-1 result at $0.52 a task plus Search grounding rather than on GPQA. This is the next crown we re-test: Claude Fable 5 tops all three boards that measure how often a model is simply wrong, at 40 on AA-Omniscience, 61% on Omniscience Accuracy and 53.3% on Humanity’s Last Exam.
Best AI for Problem Solving
Best AI for Problem Solving: GPT-5.6 Sol (#1 on LiveBench Mathematics, Reasoning and ARC-AGI-2)
The best AI for hard problem solving is GPT-5.6 Sol, and it is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 93%, the closest any model has come to the 100% human panel. Three separate houses put it first on the reasoning tasks that matter, which is more agreement than any other category on this page produces.
OpenAI has still not published Sol’s FrontierMath score, so the verified OpenAI mark remains GPT-5.5 Pro’s 39.6% on FrontierMath Tier 4, and we will slot Sol’s number in the moment it goes public. Qwen 3.7 Max is the value alternative for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, at a fraction of the cost of ChatGPT Pro, and it now includes 200 free model requests per day.
Claude Opus 5 is the alternative for long agentic reasoning chains, leading Artificial Analysis’s Agentic Index at 55.3 and ARC Prize’s ARC-AGI-3 at 30%, roughly 3.75x the next-best model. It runs second to Sol on LiveBench Reasoning at 91.2. For multimodal reasoning where the problem includes diagrams or documents, Gemini 3.1 Pro is still the practical pick.
Model | Best For | Key Benchmark | Weakness | Price |
|---|
GPT-5.6 Sol | Hardest math, science and reasoning | #1 LiveBench Mathematics (96.2), #1 LiveBench Reasoning (91.7), #1 ARC-AGI-2 (93%) | FrontierMath still unpublished; scheming flagged by METR | $100/mo ChatGPT Pro; API $5 / $30 |
Claude Opus 5 | Long agentic reasoning chains | #1 Agentic Index (55.3), #1 ARC-AGI-3 (30%) | Second on LiveBench Reasoning; 50% hallucination rate | $5 / $25 |
GPT-5.5 Pro | Verified FrontierMath leader | 39.6% FrontierMath Tier 4 | Superseded by Sol as flagship | $100/mo ChatGPT Pro |
Qwen 3.7 Max | Competition math on a budget | 97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day | API-only | $1.25 / $3.75 promo; $2.50 / $7.50 list |
Claude Fable 5 | Math inside a coding workflow | #1 on Arena’s math subcategory (1543), 96.0 LiveBench Mathematics | Priciest option here | $10 / $50 |
GLM-5.2 | Open-weight problem solving | Highest open Intelligence Index at 51, MIT, 1M context | Self-host or provider only | Open weights (MIT) |
Runner-up and alternatives: Claude Opus 5 is the runner-up and the natural pick for long-chain agentic reasoning, Claude Fable 5 is the runner-up on Arena’s math board, Qwen 3.7 Max is the value pick, and GLM-5.2 is the open-weight pick.
What changed this month: GPT-5.6 Sol keeps this crown, and it picked up a second argument. Sol now also leads Artificial Analysis’s Terminal-Bench v2.1 at 89.5% and its Coding Index at 78.3, and it matches Gemini 3.1 Pro on GPQA Diamond at 94.1%. Claude Opus 5 stays the agentic-reasoning alternative on the strength of ARC-AGI-3.
Best AI Agent
Best AI Agent: Gemini Spark vs Claude Cowork ($99.99/month Ultra vs $20/month Pro)
The best AI agent right now is Gemini Spark for 24/7 cloud-resident work and Claude Cowork for desktop-resident work, with ChatGPT Codex as the alternative for coding agents and OpenAI Operator-class browser agents as the alternative for web tasks. AI agents are the fastest-moving category of 2026: each top vendor now ships an agent product, and the practical choice is between agents that live in the cloud (run while your laptop is closed) and agents that live on your desktop (drive your apps directly).
Gemini Spark launched at Google I/O on May 19, 2026 and is the first 24/7 cloud agent. Claude Cowork launched in general availability on April 9, 2026 and runs as a desktop agent that drives your local apps. ChatGPT Codex Mobile (May 14) is the pick for coding-agent work, now usable from iOS and Android. Read the full Gemini Spark vs Claude Cowork comparison.
Agent | Best For | Where It Runs | Strength | Price |
|---|
Gemini Spark | 24/7 cloud tasks, Workspace workflows | Google Cloud VM (always-on) | First true 24/7 agent, deep Workspace integration | $99.99/mo Google AI Ultra |
Claude Cowork | Desktop, app-driving, design + code | Your Mac/Windows desktop | Drives local apps, sees your screen | $20/mo Claude Pro |
ChatGPT Codex Mobile | Coding agent on phone | OpenAI cloud + iOS/Android | Approve diffs and redirect work from phone | Included in ChatGPT plans |
Grok Agentic (Grok 4.5) | Real-time research, X scraping | xAI cloud | Native X integration | $30/mo SuperGrok |
OpenAI Operator-class | Browser tasks, web forms | OpenAI cloud + your browser | Web automation | ChatGPT Pro |
Runner-up and alternatives: Claude Cowork is the runner-up overall and the natural pick when you want the agent on your machine driving your apps. ChatGPT Codex Mobile is the runner-up for coding agents. Grok Agentic is the niche pick for real-time research.
What changed this month: no new consumer agents shipped, so the Gemini Spark (cloud) versus Claude Cowork (desktop) choice still drives most agent decisions for individual users. The model layer underneath them moved again, and Arena’s Agent board now has no open model in first place, which was true as recently as last month. Claude Opus 5 (July 24) took #1 on Artificial Analysis’s Agentic Index at 55.3, ahead of GPT-5.6 Sol at 54.0 and Claude Fable 5 at 52.8, and it costs $5 / $25. Claude Opus 5 also tops Arena’s Agent board, where the metric is task success rate rather than Elo, at 0.176 ahead of its own High tier and Kimi K3 in third at 0.142. For teams building their own agents, Meta’s Muse Spark 1.1 is a cheap agent-native option at $1.25 / $4.25 that leads Scale SEAL’s MCP Atlas tool-use board at 88.1, and GLM-5.2 (MIT) is the strongest open-weight agent model you can host on modest hardware, at #5 on Arena’s Agent board.