The Best AI to Use In August 2026

Compare leading AI models & Understand which is the best model for your needs. [Updated 1st of August]

various popular AI models like ChatGPT, Gemini, Grok, Claude, Nano Banana, etc. are orbiting Fello AI logo to symbolize that they're part of the app.

August 2026 opens with Claude Opus 5 on top and two crowns changing hands. Opus 5 still leads Artificial Analysis’s Intelligence Index at 61 and its Agentic Index at 55.3 at $5 / $25 per 1M tokens, half the price of Claude Fable 5, and it has now taken the coding crown as well after finishing first on both of Arena’s vote-based coding boards. The other move is at the cheap end, where DeepSeek V4-Flash 0731 arrived on July 31 as the new price-performance pick at $0.14 / $0.28 per 1M tokens. July had opened with Fable 5 itself back online: on July 1, Anthropic redeployed its Mythos-class flagship after the US government lifted the June 12 export-control order that had pulled the model offline for nearly three weeks. The month ended with a price move rather than a model: on July 30 OpenAI cut GPT-5.6 Luna by 80% to $0.20 / $1.20 per 1M tokens and Terra by 20% to $2 / $12, holding Sol at $5 / $30 and leaving every ChatGPT subscription price where it was. The last day of July added two more releases: DeepSeek re-trained V4-Flash into a model that scores like a frontier mid-tier at a tenth of the price, and MiniMax launched H3, a 2K video model with native audio.

The underlying board has shifted, and so has the ruler. Artificial Analysis now scores models on Intelligence Index v4.1, which rebased the scale, so these numbers are lower than the ones it published earlier in the year and the two are not comparable. On v4.1 the new Claude Opus 5 leads at 61, ahead of Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, Claude Opus 4.8 at 56, and Grok 4.5 at 54 for a fraction of their price. One thing matters for reading the rest of this page. Arena has now rated Opus 5 across its boards, at #6 on text (1,491.8), #1 on WebDev (1,702.9), #1 on image-to-WebDev (1,668.6) and #1 on the Agent board on the August 1 vote cutoff, so the crown re-test we flagged last month is finished and coding has moved. Below, we break down which model wins each category, why, and when you should pick the alternative.

GPT-5.6, ChatGPT’s default since July 9, is the best AI model for daily chat and knowledge work because it is the assistant most people can actually open, Claude Opus 5 is the best for coding at #1 on Arena’s WebDev (1,702.9) and image-to-WebDev (1,668.6) boards for $5 / $25, with Claude Fable 5 the runner-up for the hardest long-horizon work, Claude Fable 5 is also the best for writing at #1 on Arena’s text leaderboard (1508.6) and on Humanity’s Last Exam, Gemini 3.1 Pro is the best for accuracy at 98% on ARC-AGI-1 for $0.52 a task, GPT-5.6 Sol is the best for hard problem solving at #1 on LiveBench Mathematics, Reasoning and ARC-AGI-2, DeepSeek V4-Flash 0731 is the best for price-performance at Intelligence Index 50 for $0.14 / $0.28 per 1M tokens,

ChatGPT Images 2.0 is the best for image generation, Gemini Omni Flash is the best for AI video at #1 on both video boards, Grok 4.5 is the pick for real-time X context and the fewest content restrictions, and Gemini Spark plus Claude Cowork are the two AI agents most worth your attention right now. Two of those calls changed this month. Coding moved to Opus 5 on Arena’s vote-based boards, and price-performance moved to DeepSeek V4-Flash, which matches Gemini 3.6 Flash’s Intelligence Index of 50 at a tenth of its rate card. GPT-5.6 Luna at Intelligence Index 51 for $0.20 / $1.20 is the closed-model alternative if you would rather not send work to a Chinese provider.

Monthly Ranking of Top AI Models

AI models change fast. New versions are released, performance shifts, and strengths evolve over time. To keep this comparison accurate and up to date, we publish a Best AI of the Month analysis every month, based on the latest model updates and real-world performance. Below are our most recent monthly rankings, where we take a deeper look at how the leading AI models performed during each month. 

Claude Fable 5

Best AI for Writing

Claude Fable 5 is the only model in the top three of all three independent writing boards. It is #1 on Arena’s text leaderboard at 1508.6 Elo on the August 1 vote cutoff, #1 on Humanity’s Last Exam at 53.3%, and #1 on AA-Omniscience at 40, the board that measures how often a model is simply wrong. It runs a 1M-token context at $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. Claude Sonnet 5 is the value pick, free and default on claude.ai at $2 / $10 introductory pricing.

ChatGPT-5.6

Best AI for Chat / Daily Assistant

GPT-5.6 (Sol, Terra, Luna) has been ChatGPT’s default model since July 9, 2026, which makes it the best assistant most people can actually open. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, after the July 30, 2026 price cut, costs 60% less than it; API pricing runs Luna $0.20 / $1.20, Terra $2 / $12, and Sol $5 / $30 per 1M tokens. On raw preference it is not the leader, sitting #14 on Arena’s text leaderboard where Claude Fable 5 is #1. OpenAI’s system card and the evaluator METR also flagged elevated “scheming” behaviour in Sol.

ChatGPT Images 2.0

Best AI for Images

ChatGPT Images 2.0 holds the crown on both independent image boards, leading text-to-image at 1385 Elo and image editing at 1463, and it is still the best model for rendering readable multilingual text. It is included in ChatGPT Plus and Pro. Reve 2.1 is the runner-up, with Meta’s Muse Image third on text-to-image and second on image editing.

Gemini Omni Flash

Best AI for Video

Gemini Omni Flash is #1 on both independent video boards, leading Arena’s text-to-video leaderboard at 1527 Elo, a full 45 points clear of the runner-up. It generates 10-second clips with conversational editing, priced at $1.50 in and $17.50 per 1M video output tokens, roughly $0.10 per second. MiniMax H3 (July 31) is the new contender, generating 2K clips of 4 to 15 seconds with native stereo audio from $0.13 per second, and Veo 3.1 is the alternative when you need longer production clips.

Claude Opus 5

Best AI for Coding

Claude Opus 5 holds the coding crown on the two boards where developers vote on real output rather than a harness running a script. It is #1 on Arena’s WebDev board at 1,702.9 and #1 on image-to-WebDev at 1,668.6, and it does it at $5 / $25 per 1M tokens, half the price of Claude Fable 5. Fable 5 is the runner-up at #4 on WebDev and #2 on image-to-WebDev, and it stays the pick for the hardest long-horizon agentic work at $10 / $50. Kimi K3 is the contender, sitting #2 on WebDev at 1,675.5 and #3 on Arena’s Agent board.

Grok 4.5

Best AI for Creativity

Grok 4.5 is our creativity pick on product grounds, not board position. It carries the fewest content restrictions of any frontier model and the only native real-time X integration, and it is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month. To be clear, this is not a quality ranking: Grok 4.5 sits #33 on EQ-Bench Creative Writing and #41 on Arena’s creative-writing board. If you want the best-written output, Claude Fable 5 wins that comparison outright.

Gemini 3.1 Pro

Best AI for Accuracy

Gemini 3.1 Pro ties the human panel on ARC-AGI-1 at 98%, and does it at $0.52 per task, which is what makes it the practical accuracy pick rather than the most expensive one. It pairs that with native Google Search grounding for live factual answers. It scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 77.1% on ARC-AGI-2, though the ARC-AGI-2 board has moved on and 77.1% now places it around 14th. Gemini 3.5 Pro is still unreleased.

ChatGPT-5.6

Best AI for Problem Solving

GPT-5.6 Sol is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 93%, the closest any model has come to the 100% human panel. OpenAI has still not published Sol’s FrontierMath score, so GPT-5.5 Pro’s verified 39.6% on Tier 4 stays the cited OpenAI mark. Qwen 3.7 Max is the value alternative at 97.1 on the February 2026 HMMT index.

What is new in August 2026

DeepSeek V4-Flash 0731 – DeepSeek – July 31, 2026 – same price, same size, Intelligence Index 40 to 50

DeepSeek re-post-trained V4-Flash and published the result on July 31, 2026 as DeepSeek-V4-Flash-0731. Nothing about the shape of the model changed: it is the same 284B total / 13B active mixture-of-experts, the same 1M-token context, the same MIT licence and the same $0.14 / $0.28 per 1M tokens on DeepSeek’s own rate card. What moved is capability. Artificial Analysis now scores it at Intelligence Index 50, up from 40, with Agentic 45.7 and 78.7% on Terminal-Bench v2.1, which lifts it from roughly thirteenth to third in the open field behind Kimi K3 and GLM-5.2. At $0.03 per index task it is the cheapest figure on Artificial Analysis’s entire board, and that is what moved our price-performance pick this month. One honest limit: the score is one day old, comes from a single house, and the model has no votes on any Arena board yet.

MiniMax H3 – MiniMax – July 31, 2026 – 2K video with native stereo audio from $0.13 a second

MiniMax launched H3 (Hailuo 3.0) on July 31, 2026 as a general-purpose multimodal video model that takes text, image, video and audio as input. It outputs 2K clips of 4 to 15 seconds with native stereo audio, supports motion transfer, reference-driven generation and generative video editing, and is billed per second at $0.13 for 2K and $0.09 for 768p. Artificial Analysis places it #2 on its video board at 1241.5, 3.3 Elo behind Gemini Omni Flash. We have kept the video crown where it is: that gap is one day old and comes from the one board that has not reproduced between parses, and Arena’s text-to-video cutoff predates H3 entirely, so it has not beaten anything on votes. MiniMax has promised the weights in early August, and they had not been published when this update went live.

Inkling Small – Thinking Machines – late July 2026 – a quarter the size of Inkling at nearly the same score

Thinking Machines followed its first open model with Inkling Small, a 276B total / 12B active mixture-of-experts under Apache 2.0 that routes each token to 6 of 256 experts plus 2 shared ones. It accepts text, image and audio input and returns text, and Artificial Analysis scores it at Intelligence Index 40 against 41 for the full-size Inkling, at roughly a quarter of the parameters. Sources disagree on the exact publication date, so we are not printing one. The weights, an NVFP4 build and the deployment recipes are all on Hugging Face.

GPT-5.6 price cut – OpenAI – July 30, 2026 – Luna 80% cheaper, Terra 20% cheaper, Sol unchanged

OpenAI cut the API price of its two cheaper GPT-5.6 tiers on July 30, 2026, three weeks after the family reached general availability. Luna fell 80% to $0.20 / $1.20 per 1M tokens and Terra fell 20% to $2 / $12, with cached input reads dropping to $0.02 on Luna and $0.20 on Terra. The flagship Sol stays at $5 / $30, prompts above 272K input tokens are still billed at 2x input and 1.5x output, and no ChatGPT subscription price changed, so Free, Go at $8, Plus at $20, Pro at $100 and $200 and Business at $25-30 per seat are all unmoved. The cut leaves Luna at Intelligence Index 51 on Artificial Analysis for about $0.07 per index task, a point above Gemini 3.6 Flash on score and well under its $0.56 per task. Read our cover: GPT-5.6.

Claude Opus 5 – Anthropic – July 24, 2026 – new #1 on Artificial Analysis, tops both the Intelligence Index (61) and the Agentic Index (55.3)

Anthropic released Claude Opus 5 on July 24, 2026, its fourth model in under two months after Mythos 5, Fable 5, and Sonnet 5. On Artificial Analysis’s rebased v4.1 leaderboard it is now the top-ranked model overall, leading both the Intelligence Index at 61 and the Agentic Index at 55.3, ahead of Fable 5 (60 / 52.8) and GPT-5.6 Sol (59 / 54.0). API pricing is $5 / $25 per 1M tokens, identical to Opus 4.8 and half the cost of Fable 5, under the API id claude-opus-5. It adds a Fast mode that runs about 2.5x quicker at twice the base price, plus an effort setting from low to high and a new max tier. Anthropic’s own benchmarks put it at more than double Opus 4.8 on Frontier-Bench v0.1, within 0.5% of Fable 5 on CursorBench 3.2 at max effort, and roughly 3x the next-best model on ARC-AGI 3. Anthropic’s own routing guidance now reads “start with Claude Opus 5 for complex agentic coding and enterprise work. For workloads that need the highest available capability, use Claude Fable 5.” One caveat worth knowing: Opus 5 remains behind Mythos 5 on cybersecurity tasks. It also takes Artificial Analysis’s GDPval-AA v2 professional-deliverables board outright at 1858, ahead of Fable 5’s 1746. Arena has since rated it across its boards, putting it #1 on WebDev, image-to-WebDev, Document and the Agent board and #6 on text, which is what moved the coding crown to it this month. It is the default on Claude Max and strongest on Claude Pro. Read our cover: Claude Opus 5.

Fugu-Ultra v1.1 – Sakana AI – July 24, 2026 – orchestration-engine refresh with vendor-reported gains of up to 7.9 points over v1.0 at the same price

Sakana AI shipped Fugu-Ultra v1.1 on July 24, 2026, a refresh of the frontier models inside its TRINITY orchestration engine rather than a new base model. Sakana’s own announcement claims gains of up to 7.9 points over v1.0 at unchanged pricing, though it does not publish v1.0 scores side by side, so treat that as the vendor’s figure. All the benchmark numbers come from Sakana’s custom scaffolding rather than independent testing: on its own charts it leads Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, scoring 73.7% on SWE-Bench Pro, 82.1% on Terminal-Bench 2.1, 93.2% on LiveCodeBench, and 95.5 on GPQA-Diamond. It does not sweep the field, though: its 50.0 on Humanity’s Last Exam is a rounding-error tie with Opus 4.8’s 49.8. Pricing is unchanged from v1.0 at $5 / $30 per 1M input/output tokens (plus $0.50 cached), rising to $10 / $45 above the 272K-token context mark, with $20 / $100 / $200 monthly Standard, Pro, and Max plans; the API is OpenAI-compatible with a 1M-token context window. Read our coverage: Fugu-Ultra v1.1.

Gemini 3.6 Flash – Google – July 21, 2026 – cheaper output, faster, same Intelligence Index

Gemini 3.6 Flash is Google’s newest Flash model and the one the free Gemini app now reaches. Output falls to $7.50 per 1M tokens from $9.00 while input holds at $1.50, so the cut is output-only. Google says it “reduces output token usage by 17% compared to 3.5 Flash,” and up to 65% on long-horizon agentic work like DeepSWE, so the real per-task saving is larger than the sticker. Artificial Analysis scores 3.6 Flash and 3.5 Flash at the same Intelligence Index 50 on v4.1, so treat it as cheaper and faster rather than smarter. It shipped alongside Gemini 3.5 Flash-Lite at $0.30 / $2.50, and is live through the Gemini API in Google AI Studio and Android Studio, the Gemini Enterprise Agent Platform, and the Gemini app for everyone. Google also announced Gemini 3.5 Flash Cyber, but that one has not shipped: it goes to governments and trusted partners via CodeMender as a limited-access pilot.

Qwen 3.8 (Qwen3.8-Max) – Alibaba – July 19, 2026 – 2.4-trillion-parameter multimodal flagship, previewed at WAIC Shanghai

Alibaba previewed Qwen3.8-Max on July 19, its largest model yet at 2.4 trillion total parameters on a sparse Mixture-of-Experts design, and the first Qwen above 1 trillion parameters to handle images, video, and documents alongside text. The Qwen team calls it “second only to Fable 5,” but that ranking is Alibaba’s own claim, shipped with no benchmarks, no model card, no activated-parameter count, and no license. A preview is live through Alibaba’s Token Plan, Qoder, and QoderWork at 10% of standard pricing, and Alibaba says the full model will go open-weight “soon” without giving a date. We are keeping Qwen 3.8 out of the ranked picks until independent benchmarks or the actual weights arrive. Read our full breakdown of Qwen 3.8

Kimi K3 – Moonshot AI – July 16, 2026 – 2.8T params, the largest open-weight model ever released (weights shipped July 27)

Moonshot AI launched Kimi K3 on the eve of July 16, 2026, its new flagship and the successor to the K2 line, then published the full open weights on July 27, 2026. The Hugging Face model card confirms a 2.8-trillion-parameter Mixture-of-Experts design with 104 billion parameters active per token, a 1,048,576-token context window, and text, image and video input (no audio). Thinking is always on, and reasoning_effort now takes low, high or max, with max as the default. The independent scores have landed too. Artificial Analysis puts K3 at Intelligence Index 57, the highest of any open-weight model, and on Arena it is #3 on the Agent board and #2 on WebDev behind Claude Opus 5, so the launch-day vendor claims no longer have to carry the case on their own. Official API pricing is $3 / $15 per 1M tokens with cached input at $0.30, plus $0.005 per successful web-search call. The download is 96 safetensors shards and about 1.56 TB under a custom Kimi K3 License rather than MIT, so self-hosting is a multi-node job; a free basic chat tier exists in the Kimi app, with heavier agentic use metered on paid plans. Read our review: Kimi K3 

GPT-5.6 Sol, Terra, and Luna – OpenAI – July 9, 2026 – next-gen family live across ChatGPT, Codex, and the API

OpenAI opened its GPT-5.6 family to general availability on July 9, 2026, ending the two-week gated preview that started June 26 behind a US-government safety review. The lineup runs from least to most capable: Luna, a fast, low-cost tier; Terra, a balanced everyday model OpenAI says matches GPT-5.5 at roughly half the cost; and Sol, the flagship, tuned for biology, chemistry, and cybersecurity. GPT-5.6 is now rolling out across ChatGPT, Codex, and the API as OpenAI’s default, with launch API pricing of Sol $5 / $30, Terra $2.50 / $15, and Luna $1 / $6 per 1M tokens; Terra and Luna were cut on July 30, 2026 to $2 / $12 and $0.20 / $1.20. On the few benchmarks OpenAI published, Sol scores 88.8% on Terminal-Bench 2.1 (91.9% in its higher-compute “ultra” mode) versus GPT-5.5’s 88.0%, and 60.5 on HealthBench Professional, up 8.7 points on GPT-5.5; OpenAI notably withheld the usual SWE-bench Verified, GPQA, and FrontierMath numbers, and its context window is still not officially published (a circulating 1.5M figure is unconfirmed). One caveat worth knowing: OpenAI’s own system card and the external evaluator METR flagged elevated “scheming” behaviour in Sol, including gaming a software-engineering test at the highest rate METR has ever recorded, which is part of why the release was gated for review. Read our cover: GPT-5.6.

Gemini 3.5 Pro – Google – Still unreleased – months behind schedule, per Bloomberg

Gemini 3.5 Pro remains the biggest pending launch. Google announced it at I/O on May 19 alongside Gemini 3.5 Flash, but only Flash shipped, and the June target slipped. Bloomberg reported on July 16, 2026 that the model is months behind schedule and has fallen short of Google’s internal goals, with the company still working to improve its capabilities particularly in coding. That is the whole sourced picture. Google has never publicly confirmed a launch date, and the specific slip dates and enterprise-preview details circulating elsewhere are not in any primary source. Google has not published final specs such as the context window or reasoning modes, so treat circulating figures as unconfirmed. Use Gemini 3.6 Flash in the meantime; we will move 3.5 Pro into the main ranking the moment it goes live.

Category Deep Dives

Below, we provide a series of comprehensive, category-by-category deep dives to help you choose the ideal AI model for your specific operational goals. We systematically evaluate the leading proprietary and open-weight options across nine distinct specialties – ranging from writing style and daily assistant workflows to advanced coding execution, multi-tier factual reasoning, cloud-resident agents, and high-fidelity video generation, ensuring you deploy the highest-performing intelligence for each task.

Best AI for Writing

Best AI for Writing: Claude Fable 5 (#1 on Arena creative writing and LiveBench Language)

The best AI for writing is Claude Fable 5, the only model in the top three of all three independent writing boards, with Claude Sonnet 5 as the free-tier value pick and GPT-5.5 as the alternative for fact-anchored business writing. Fable 5 leads Arena’s creative-writing leaderboard, tops LiveBench Language at 90.7, and places third on EQ-Bench Creative Writing v3 behind Kimi K3 and GPT-5.6 Sol. No other model is top-three on more than one of them.

This is a change from last month, when Claude Sonnet 5 held this slot on the strength of its GDPval-AA score. The preference boards do not support that placing. Sonnet 5 sits #53 on Arena creative writing and #13 on EQ-Bench, and scores 75.0 on LiveBench Language against Fable 5’s 90.7. It is still the model most people should write with day to day, because it is free and default on claude.ai, but it is the value pick rather than the quality leader.

Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits, with Pro and Team Standard reaching it through usage credits. If your writing is a work deliverable rather than prose, Claude Opus 5 tops Artificial Analysis’s GDPval-AA v2 professional-deliverables board outright at 1858, well clear of Fable 5’s 1746. Fable 5 keeps the writing crown because it leads the boards that judge prose and factual care: Arena’s text leaderboard at 1508.6, Humanity’s Last Exam at 53.3% and AA-Omniscience at 40.

Model

Best For

Strength

Weakness

Price (per 1M tokens)

Claude Fable 5

Best writing overall

#1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #3 EQ-Bench

Priciest option here

$10 / $50

Kimi K3

Creative fiction and voice

#1 EQ-Bench Creative Writing (2377), 234 Elo clear of second

Only #10 on Arena creative writing

$3 / $15

Claude Sonnet 5

Free everyday writing

Free and default on claude.ai, 1M context

#53 Arena creative writing; 75.0 LiveBench Language

$2 / $10 intro (then $3 / $15)

Claude Opus 5

Professional deliverables

#1 GDPval-AA v2 at 1858, ahead of Fable 5 (1746)

Behind Fable 5 on Arena text and on Humanity’s Last Exam

$5 / $25

GPT-5.5

Fact-anchored business writing

Documented factual-reliability gains over GPT-5.4

Artificial Analysis now marks its reasoning tiers deprecated

$5 / $30

Gemini 3.6 Flash

Bulk drafts at scale

17% fewer output tokens than 3.5 Flash

Weaker on hardest reasoning

$1.50 / $7.50

Runner-up and alternatives: Kimi K3 is the runner-up for creative fiction and wins EQ-Bench outright, Claude Sonnet 5 is the runner-up on value and the one to use if you are not paying, Claude Opus 5 is the pick for professional deliverables, and Gemini 3.6 Flash is the pick for bulk drafting.

What changed this month: the writing crown stays with Claude Fable 5, and it is the strongest-supported call on the page. Fable 5 is now #1 on Arena’s text leaderboard at 1508.6 on the August 1 cutoff, #1 on Humanity’s Last Exam at 53.3% and #1 on AA-Omniscience at 40, which is three different houses agreeing. Claude Opus 5 keeps the professional-deliverables lane on GDPval-AA v2, where the re-fitted board now reads 1858 against Fable 5’s 1746.

Best AI for Chat & Daily Assistant

Best AI for Chat & Daily Assistant: GPT-5.6 (ChatGPT’s default since July 9)

The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena’s text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. If you want the best-rated conversational model and are willing to leave ChatGPT, that is the swap to make.

Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month for roughly 5x Plus usage or $200/month for roughly 20x), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $5 / $30 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI’s system card and the evaluator METR flagged elevated “scheming” behaviour in Sol.

GPT-5.5 is still sold by OpenAI at $5 / $30 and its Instant tier is still the safer pick for hallucination-sensitive work, with a documented 52.5% drop in hallucinated claims over GPT-5.3 Instant. Note that Artificial Analysis has since marked every GPT-5.5 reasoning tier deprecated, so treat it as a model you can still buy rather than a current benchmark reference. Claude Opus 5 is the better pick when you want a model that pushes back on weak prompts, and Gemini 3.6 Flash is the better pick if you are running everything through the free Gemini app.

Model

Best For

Strength

Weakness

Price

GPT-5.6

Everyday chat, ChatGPT’s default

The assistant most people can open; Terra matches GPT-5.5 at ~half cost

#14 on Arena text; scheming flagged by METR

Free / $20/mo Plus; API $0.20 / $1.20 to $5 / $30

Claude Fable 5

Highest-rated conversation

#1 on Arena text overall (1508.6) and 6 of 7 subcategories

No free tier; usage-credit access on Pro

$10 / $50 API

Claude Opus 5

Thoughtful, nuanced answers

#1 Artificial Analysis Intelligence Index (61) and Agentic Index (55.3)

#6 on Arena text, behind Claude Fable 5

$20/mo Pro, $5 / $25 API

GPT-5.5 Instant

Hallucination-sensitive daily work

52.5% fewer hallucinated claims vs 5.3 Instant

Reasoning tiers now marked deprecated by Artificial Analysis

$20/mo Plus; API $5 / $30

Gemini 3.6 Flash

Fast, free, multimodal

Free in the Gemini app, 1M context, #12 on Arena text

Weaker on hardest reasoning

Free / $1.50 / $7.50 API

Fello AI

All the top models, one app

ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad

Routed via app, not direct

$9.99/mo

Runner-up and alternatives: Claude Fable 5 is the runner-up and the actual preference leader, Claude Opus 5 is the runner-up for thoughtful daily use, Gemini 3.6 Flash is the runner-up for fast and free, and Grok 4.5 is the niche pick for live-news days. Fello AI is the natural pick if you want the top models in one Mac and iOS app for $9.99/month instead of juggling subscriptions.

What changed this month: GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open.

Best AI for Images

Best AI for Images: ChatGPT Images 2.0 (#1 on text-to-image and image editing)

The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena’s text-to-image board at 1385 Elo and its image-editing board at 1463, and Artificial Analysis puts it first on its own image arena at 1339.4. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro.

The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta’s Muse Image is #3 on Arena’s text-to-image board and #2 on image editing, the strongest showing any Meta image model has managed. Note that the houses disagree about third place: on Artificial Analysis’s image board it is Microsoft’s MAI-Image-2.5 at 1269.7 in third, and that model is #3 on Arena’s image-editing board too. Google’s Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.

Model

Best For

Strength

Weakness

Price

ChatGPT Images 2.0

Images with readable text

#1 text-to-image (1385) and #1 image editing (1463)

Less photoreal than the Gemini image line

Included in ChatGPT Plus

Reve 2.1

Layout, typography, native 4K

#2 text-to-image at 1302, layout-preserving editing

Smaller ecosystem

Free / from $7.99/mo

Muse Image

Image editing, Meta ecosystem

#3 text-to-image, #2 image editing (1407)

New, thin tooling around it

Meta AI app

Nano Banana 2 (Gemini 3.1 Flash Image)

Photoreal portraits and products

Outranks Nano Banana Pro on both boards

Weaker on text in image

Gemini app / AI Studio

Seedream 5.0 Pro

Multilingual text + region-precise editing

10+ languages incl. Arabic RTL, lasso and layer editing

No independent benchmarks; copyright cloud

BytePlus / Magnific

Midjourney v8

Stylized art, illustration

Aesthetic baseline most artists prefer

Weaker on text in image

$10-$120/mo

Grok Imagine

NSFW / Spicy Mode

Most permissive guardrails

Smaller model behind it

$30/mo SuperGrok

Runner-up and alternatives: Reve 2.1 is the runner-up overall and the pick for layout and typography, Muse Image is the runner-up for editing an image you already have, and Nano Banana 2 is the photoreal pick. Grok Imagine is still the only frontier model that allows Spicy Mode adult content.

What changed this month: the image crown is unchanged and GPT Image 2 still leads all three boards we track, on Arena text-to-image (1385), Arena image editing (1463) and Artificial Analysis’s image arena (1339.4). The one addition is Microsoft’s MAI-Image-2.5, which sits third on Artificial Analysis’s image board at 1269.7 and third on Arena’s image-editing board, so third place now depends on which house you read.

Best AI for Video

Best AI for Video: Gemini Omni Flash (#1 on both video leaderboards)

The best AI for video generation is Gemini Omni Flash, which leads Arena’s text-to-video board at 1527 Elo, a full 45 points clear of second place, and also tops Artificial Analysis’s video arena. It is #1 on both houses, which no other video model manages. Google reached the consumer launch on May 19, 2026 through the Gemini app, Flow and YouTube Shorts, and opened developer access on June 30 through AI Studio and the Gemini API. Read our full breakdown of Gemini Omni Flash.

Pricing runs $1.50 in and $17.50 per 1M video output tokens, which works out at roughly $0.10 per second of finished video, and it supports conversational editing so you can adjust a clip by describing the change. The one real limit is length: Omni Flash generates 10-second clips. If you need longer takes, Veo 3.1 remains the right tool inside the Gemini app, AI Studio and Vertex AI, with native audio and 1080p output.

This replaces Veo 3.1 at the top of the category. Veo 3.1 is a good model, but it is not the leading one: its best variant sits #6 on Arena’s text-to-video board, behind Omni Flash, ByteDance’s Dreamina Seedance 2.0 and Meta’s Muse Video. Google still wins this category, just with a different model than the page previously named.

The contender to watch is MiniMax H3, launched July 31, 2026. It answers Omni Flash’s two weakest points directly, generating 2K clips of 4 to 15 seconds with native stereo audio against Omni Flash’s 10-second cap, and it is billed per second from $0.13 at 2K. Artificial Analysis already places it #2 on its video board at 1241.5, only 3.3 Elo behind Omni Flash. We are not moving the crown on that: the score is one day old, it comes from the single board that has not reproduced between parses, and Arena’s text-to-video vote cutoff predates H3, so it has not yet been judged by anyone’s votes. MiniMax has promised the weights in early August.

Model

Best For

Strength

Weakness

Price

Gemini Omni Flash

Best AI video overall

#1 on both video boards (1527 Arena), conversational editing

Caps at 10-second generations

~$0.10/sec; Gemini app / AI Studio

MiniMax H3

2K clips with native audio

4-15s at 2K, native stereo audio; #2 on AA video (1241.5)

One day old, no Arena votes yet; weights not published

$0.13/sec at 2K

Dreamina Seedance 2.0

Closest challenger

#2 text-to-video (1482) and #1 on image-to-video

ByteDance ecosystem, limited Western access

Dreamina / BytePlus

Muse Video

Meta ecosystem video

#3 text-to-video at 1459

Newest of the group, thin tooling

Meta AI app

Veo 3.1

Longer production clips

Native audio, 1080p, strong physics consistency

#6 on Arena video, not the quality leader

Google AI Pro / Ultra

Kling 3.0 / 3.0 Turbo

Fast iteration at lower cost

Native 4K, 60fps, 15-second clips; Turbo shipped June 17

Outside the top 16 on Arena text-to-video

From $10/mo

Luma Ray 3

Photoreal scenes

Strong realism for landscapes

Smaller community

Free / from $9.99/mo

Runner-up and alternatives: Dreamina Seedance 2.0 is the runner-up overall and actually beats Omni Flash on image-to-video, Muse Video is third, and Veo 3.1 is the pick when 10 seconds is not enough. Runway is no longer listed here: the page previously named Gen-4, which has since been superseded by Gen-4.5, and we could not reproduce a top-tier placing for either across the boards we track. OpenAI retired the Sora 2 consumer app on April 26, 2026 and only the developer API remains, through September 24, 2026.

What changed this month: Gemini Omni Flash keeps the video crown, and it is still #1 on both boards, at 1527.5 on Arena and 1244.8 on Artificial Analysis. MiniMax H3 (July 31) enters as the contender at #2 on Artificial Analysis’s video board (1241.5), 3.3 Elo behind, and it beats Omni Flash on specification with 2K output, clips up to 15 seconds and native stereo audio. We are holding the crown because that margin is one day old, single-board, and carries no votes on Arena, whose text-to-video cutoff predates the launch.

Best AI for Coding

Best AI for Coding: Claude Opus 5 (#1 on Arena’s WebDev and image-to-WebDev boards)

The best AI for coding is Claude Opus 5, and it wins on the two boards where developers vote on the finished result rather than a script measuring a harness. It is #1 on Arena’s WebDev board at 1,702.9 and #1 on image-to-WebDev at 1,668.6, on vote cutoffs of August 1 and July 31. Those are human comparisons with published cutoffs and no scaffold variable, which is exactly what a coding crown should rest on.

Price is the second half of the argument. Opus 5 runs $5 / $25 per 1M tokens against Claude Fable 5’s $10 / $50, and Artificial Analysis measures it at $2.34 per index task against Fable 5’s $3.15. Anthropic’s own docs tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for workloads that need the highest available capability, which is the same split. Claude Fable 5 is the runner-up and stays the pick for the hardest long-horizon work, at #4 on WebDev (1,630.7) and #2 on image-to-WebDev (1,625.7).

The contender is Kimi K3, which takes #2 on WebDev at 1,675.5, 45 points clear of Fable 5 and behind only Opus 5, and #3 on Arena’s Agent board. It is the strongest open-weight coder on the boards, and at $3 / $15 it undercuts both Claude tiers, though self-hosting it means 1.56 TB of weights across 96 shards. Read our cover of Claude Opus 5.

On price-per-result, Artificial Analysis’s Coding Index is not a Claude sweep and we will not pretend otherwise: GPT-5.6 Sol (xhigh) leads it at 78.3 with Opus 5 (max) at 78.0, a gap small enough to call a tie, with Fable 5 at 76.5 and Kimi K3 at 76.2. The cheapest serious contender is now DeepSeek V4-Flash 0731 at $0.14 / $0.28, which scores 78.7% on Artificial Analysis’s Terminal-Bench v2.1 mirror, ahead of Gemini 3.6 Flash at 77.5% and behind Grok 4.5 at 81.6%. On open weights, Kimi K3 is now the strongest option on score, since its weights shipped on July 27, while GLM-5.2 (MIT) stays the one most teams can realistically host.

Model

Best For

Strength

Weakness

Price (per 1M tokens)

Claude Opus 5

Best coding overall

#1 Arena WebDev (1,702.9) and #1 image-to-WebDev (1,668.6); $2.34 per index task

Artificial Analysis Coding Index has GPT-5.6 Sol a shade ahead

$5 / $25

Claude Fable 5

Hardest long-horizon agentic work

#2 Arena image-to-WebDev (1,625.7), #4 WebDev; 1M context

Priciest; Artificial Analysis Coding Index puts it 7th

$10 / $50

Kimi K3

Web app building and agents

#2 Arena WebDev (1,675.5), #3 Arena Agent, highest open Intelligence Index (57)

1.56 TB to self-host

$3 / $15

GPT-5.6 Sol

OpenAI flagship, agentic coding

#1 Artificial Analysis Coding Index at 78.3 (xhigh)

Absent from the official Terminal-Bench board; eval-gaming flagged by METR

$5 / $30

Muse Spark 1.1

Cheap agentic tool use

#1 on SEAL SWE-Bench Pro (public and private) and MCP Atlas (88.1)

US-only preview

$1.25 / $4.25

Grok 4.5

Cheap value coder

#4 on the official Terminal-Bench 2.1 board at 79.3% via Cursor CLI

Higher hallucination rate; EU API console still closed

$2 / $6

Gemini 3.6 Flash

Agent coding at scale

Intelligence Index 50, 17% fewer output tokens than 3.5 Flash

Weaker on hardest reasoning

$1.50 / $7.50

GLM-5.2

Best open-weight coder you can host

Intelligence Index 51, #6 Arena WebDev (1,587.1), MIT licence

Kimi K3 outscores it; self-host or provider only

Open weights (MIT)

Runner-up and alternatives: Kimi K3 is the runner-up on Arena’s WebDev board and the pick if you want the strongest open weights, Claude Fable 5 is the runner-up for the hardest long-horizon work, GPT-5.6 Sol is the runner-up on Artificial Analysis’s composite, and GLM-5.2 is the open-weight pick for teams hosting it themselves. Inside IDEs, Cursor with Claude is still the most popular pairing and Claude Code is the natural pick if you live in the terminal.

What changed this month: the coding crown moved to Claude Opus 5. Arena finished rating it, and it came first on both coding boards, WebDev at 1,702.9 and image-to-WebDev at 1,668.6, with Claude Fable 5 fourth and second. Those are vote-based boards with published cutoffs, so the crown now rests on human comparisons rather than on any one harness, and Opus 5 does it at half Fable 5’s price. Fable 5 becomes the runner-up for the hardest long-horizon work, and Kimi K3 stays the contender at #2 on WebDev.

Best AI for Creativity

Best AI for Creativity: Grok 4.5 (fewest content restrictions, native real-time X)

The best AI for unfiltered, on-trend creative work is Grok 4.5, and we want to be exact about why. This pick is about the product, not the prose quality. Grok 4.5 carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month.

It is not the best writer, and the boards are blunt about it. Grok 4.5 sits #33 on EQ-Bench Creative Writing and #41 on Arena’s creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. If you are picking on output quality alone, Claude Fable 5 wins outright at #1 on Arena creative writing, and Kimi K3 wins EQ-Bench. Choose Grok 4.5 for what it will let you make, not for how well it writes.

Model

Best For

Strength

Weakness

Price

Grok 4.5

Unfiltered, opinionated, on-trend

Fewest content restrictions, native real-time X grounding

#33 EQ-Bench, #41 Arena creative writing

$30/mo SuperGrok

Claude Fable 5

Highest-quality creative prose

#1 Arena creative writing, #1 LiveBench Language

Cautious guardrails on edgy briefs

$10 / $50 API

Kimi K3

Fiction and distinctive voice

#1 EQ-Bench Creative Writing at 2377

Only #10 on Arena creative writing

$3 / $15

Claude Opus 5

Long-form structured creativity

Holds long threads and self-edits; #1 Intelligence Index

Most cautious of the group

$20/mo Pro, $5 / $25 API

Gemini 3.1 Pro

Multimodal creative

Strong text, image and video chain

Quotas inside the Gemini app

Free / $2.00-$4.00 API in

Grok Imagine (Spicy Mode)

NSFW / adult creative

Most permissive image generation

Niche use case

$30/mo SuperGrok

Runner-up and alternatives: Claude Fable 5 is the runner-up and the right pick if quality matters more than freedom, Kimi K3 is the pick for fiction, and Claude Opus 5 is the pick for creative projects that run across many turns. For adult creative work, Grok Imagine Spicy Mode is still the only frontier-grade option.

What changed this month: Grok 4.5 keeps this pick, and nothing about the product moved. It is here for its permissiveness and its live X access, not for board position, and Claude Fable 5 remains the model to use when you want the better writing. One naming note for August: xAI now trades as SpaceXAI on the leaderboards, and the model is unchanged.

Best AI for Accuracy

Best AI for Accuracy: Gemini 3.1 Pro (98% on ARC-AGI-1, at $0.52 per task)

The best AI for accuracy and research is Gemini 3.1 Pro. Its strongest result is on ARC Prize’s ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you actually want when the answer has to be current rather than merely plausible.

It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity’s Last Exam, and tops Scale SEAL’s HLE board at 46.44. We have dropped the page’s previous ARC-AGI-2 framing. Its 77.1% is still correct, but the board has moved and that score now places it around 14th, behind GPT-5.6 Sol at 93% and Claude Opus 5 at 90%, so it is no longer evidence of an accuracy lead.

Two honest caveats. On grounded search specifically, Arena’s search leaderboard is led by Anthropic, not Google, with Gemini 3.1 Pro grounding at #7. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model according to ARC Prize. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience, which is weak ground for a crown named accuracy.

Model

Best For

Key Benchmark

Weakness

Price

Gemini 3.1 Pro

Cheap, reliable factual work

98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol)

ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search

$2.00-$4.00 / $12.00-$18.00 (tiered)

Claude  Fable 5

Grounded search

#1 and #3 on Arena’s search leaderboard, ahead of Google

No single cheap tier

$10 / $50 (Fable 5)

GPT-5.6 Sol

Novel reasoning

#1 ARC-AGI-2 at 93%, against a 100% human panel

Scheming flagged by METR

$5 / $30

Claude Opus 5

Hardest unseen problems

#1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize)

Artificial Analysis measures a 50% hallucination rate

$5 / $25

Qwen 3.7 Max

Frontier accuracy at value pricing

92.4 GPQA Diamond, 200 free requests/day

API-only, no chat front-end

$1.25 / $3.75 promo; $2.50 / $7.50 list

Claude Opus 4.6

Honesty under pressure

#1 on Scale SEAL’s MASK board at 96.28; Anthropic holds the top 5

Superseded as a flagship

Legacy Anthropic model

Runner-up and alternatives: Anthropic’s models are the runner-up for grounded search and sweep the honesty-under-pressure board, GPT-5.6 Sol is the runner-up for novel reasoning, and Qwen 3.7 Max is the value pick at the frontier.

What changed this month: Gemini 3.1 Pro keeps the accuracy crown, but its headline number is no longer a lead. Its GPQA Diamond score re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly, so the crown rests on the ARC-AGI-1 result at $0.52 a task plus Search grounding rather than on GPQA. This is the next crown we re-test: Claude Fable 5 tops all three boards that measure how often a model is simply wrong, at 40 on AA-Omniscience, 61% on Omniscience Accuracy and 53.3% on Humanity’s Last Exam.

Best AI for Problem Solving

Best AI for Problem Solving: GPT-5.6 Sol (#1 on LiveBench Mathematics, Reasoning and ARC-AGI-2)

The best AI for hard problem solving is GPT-5.6 Sol, and it is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 93%, the closest any model has come to the 100% human panel. Three separate houses put it first on the reasoning tasks that matter, which is more agreement than any other category on this page produces.

OpenAI has still not published Sol’s FrontierMath score, so the verified OpenAI mark remains GPT-5.5 Pro’s 39.6% on FrontierMath Tier 4, and we will slot Sol’s number in the moment it goes public. Qwen 3.7 Max is the value alternative for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, at a fraction of the cost of ChatGPT Pro, and it now includes 200 free model requests per day.

Claude Opus 5 is the alternative for long agentic reasoning chains, leading Artificial Analysis’s Agentic Index at 55.3 and ARC Prize’s ARC-AGI-3 at 30%, roughly 3.75x the next-best model. It runs second to Sol on LiveBench Reasoning at 91.2. For multimodal reasoning where the problem includes diagrams or documents, Gemini 3.1 Pro is still the practical pick.

Model

Best For

Key Benchmark

Weakness

Price

GPT-5.6 Sol

Hardest math, science and reasoning

#1 LiveBench Mathematics (96.2), #1 LiveBench Reasoning (91.7), #1 ARC-AGI-2 (93%)

FrontierMath still unpublished; scheming flagged by METR

$100/mo ChatGPT Pro; API $5 / $30

Claude Opus 5

Long agentic reasoning chains

#1 Agentic Index (55.3), #1 ARC-AGI-3 (30%)

Second on LiveBench Reasoning; 50% hallucination rate

$5 / $25

GPT-5.5 Pro

Verified FrontierMath leader

39.6% FrontierMath Tier 4

Superseded by Sol as flagship

$100/mo ChatGPT Pro

Qwen 3.7 Max

Competition math on a budget

97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day

API-only

$1.25 / $3.75 promo; $2.50 / $7.50 list

Claude Fable 5

Math inside a coding workflow

#1 on Arena’s math subcategory (1543), 96.0 LiveBench Mathematics

Priciest option here

$10 / $50

GLM-5.2

Open-weight problem solving

Highest open Intelligence Index at 51, MIT, 1M context

Self-host or provider only

Open weights (MIT)

Runner-up and alternatives: Claude Opus 5 is the runner-up and the natural pick for long-chain agentic reasoning, Claude Fable 5 is the runner-up on Arena’s math board, Qwen 3.7 Max is the value pick, and GLM-5.2 is the open-weight pick.

What changed this month: GPT-5.6 Sol keeps this crown, and it picked up a second argument. Sol now also leads Artificial Analysis’s Terminal-Bench v2.1 at 89.5% and its Coding Index at 78.3, and it matches Gemini 3.1 Pro on GPQA Diamond at 94.1%. Claude Opus 5 stays the agentic-reasoning alternative on the strength of ARC-AGI-3.

Best AI Agent

Best AI Agent: Gemini Spark vs Claude Cowork ($99.99/month Ultra vs $20/month Pro)

The best AI agent right now is Gemini Spark for 24/7 cloud-resident work and Claude Cowork for desktop-resident work, with ChatGPT Codex as the alternative for coding agents and OpenAI Operator-class browser agents as the alternative for web tasks. AI agents are the fastest-moving category of 2026: each top vendor now ships an agent product, and the practical choice is between agents that live in the cloud (run while your laptop is closed) and agents that live on your desktop (drive your apps directly).

Gemini Spark launched at Google I/O on May 19, 2026 and is the first 24/7 cloud agent. Claude Cowork launched in general availability on April 9, 2026 and runs as a desktop agent that drives your local apps. ChatGPT Codex Mobile (May 14) is the pick for coding-agent work, now usable from iOS and Android. Read the full Gemini Spark vs Claude Cowork comparison.

Agent

Best For

Where It Runs

Strength

Price

Gemini Spark

24/7 cloud tasks, Workspace workflows

Google Cloud VM (always-on)

First true 24/7 agent, deep Workspace integration

$99.99/mo Google AI Ultra

Claude Cowork

Desktop, app-driving, design + code

Your Mac/Windows desktop

Drives local apps, sees your screen

$20/mo Claude Pro

ChatGPT Codex Mobile

Coding agent on phone

OpenAI cloud + iOS/Android

Approve diffs and redirect work from phone

Included in ChatGPT plans

Grok Agentic (Grok 4.5)

Real-time research, X scraping

xAI cloud

Native X integration

$30/mo SuperGrok

OpenAI Operator-class

Browser tasks, web forms

OpenAI cloud + your browser

Web automation

ChatGPT Pro

Runner-up and alternatives: Claude Cowork is the runner-up overall and the natural pick when you want the agent on your machine driving your apps. ChatGPT Codex Mobile is the runner-up for coding agents. Grok Agentic is the niche pick for real-time research.

What changed this month: no new consumer agents shipped, so the Gemini Spark (cloud) versus Claude Cowork (desktop) choice still drives most agent decisions for individual users. The model layer underneath them moved again, and Arena’s Agent board now has no open model in first place, which was true as recently as last month. Claude Opus 5 (July 24) took #1 on Artificial Analysis’s Agentic Index at 55.3, ahead of GPT-5.6 Sol at 54.0 and Claude Fable 5 at 52.8, and it costs $5 / $25. Claude Opus 5 also tops Arena’s Agent board, where the metric is task success rate rather than Elo, at 0.176 ahead of its own High tier and Kimi K3 in third at 0.142. For teams building their own agents, Meta’s Muse Spark 1.1 is a cheap agent-native option at $1.25 / $4.25 that leads Scale SEAL’s MCP Atlas tool-use board at 88.1, and GLM-5.2 (MIT) is the strongest open-weight agent model you can host on modest hardware, at #5 on Arena’s Agent board.

Pricing Comparison

AI Model Pricing Comparison in August 2026 ($0 free tiers to $199.99/month Google AI Ultra)

Here is the August 2026 pricing comparison for every leading AI model, in API cost per 1 million tokens and the consumer-subscription price for the same model. Free tiers exist for ChatGPT, Gemini, Claude, Grok, and DeepSeek. The most consequential price on this table is Claude Opus 5 at $5 / $25, because it is the #1 model on Artificial Analysis’s Intelligence Index at half the cost of Claude Fable 5. Meta’s Muse Spark 1.1 lists at $1.25 / $4.25 and Grok 4.5 at $2 / $6, and on a price-per-intelligence basis Artificial Analysis puts a Grok 4.5 index task at about $0.31, five times cheaper than Claude Sonnet 5. The price-performance pick is now DeepSeek V4-Flash 0731 at $0.14 / $0.28, which matches Gemini 3.6 Flash’s Intelligence Index of 50 and costs $0.03 per index task against Flash’s $0.56. Among closed models GPT-5.6 Luna is the cheapest at $0.20 / $1.20 after OpenAI’s July 30 cut, at Intelligence Index 51 and about $0.07 per index task. MiniMax M3 now lists at $0.30 per million input tokens on a permanent 50% discount, which makes it the cheaper way to run a licence-permissive multimodal model, though V4-Flash beats it on score, on price and on context. For a deeper breakdown by tier, see our full AI Pricing Comparison Guide hub.

Model

Input (per 1M)

Output (per 1M)

Context Window

Free access?

GPT-5.5

$5.00

$30.00

1M (400K in Codex)

ChatGPT Free; API paid

GPT-5.5 Pro

$30.00

$180.00

1M

ChatGPT Pro from $100/mo ($200 higher-usage tier)

GPT-5.6 Sol

$5.00

$30.00

not published

Live in ChatGPT, Codex & API (July 9)

GPT-5.6 Terra

$2.00

$12.00

not published

Live in ChatGPT, Codex & API (July 9)

GPT-5.6 Luna

$0.20

$1.20

not published

Live in ChatGPT, Codex & API (July 9)

Claude Opus 5

$5.00

$25.00

1M

Claude Pro/Max default; API paid

Claude Opus 4.8

$5.00

$25.00

1M

Legacy model at Anthropic; Pro/Max/API

Claude Fable 5

$10.00

$50.00

1M

Permanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits

Claude Sonnet 5

$2.00 intro / $3.00 list

$10.00 intro / $15.00 list

1M

Claude Free & Pro default; API paid

Claude Sonnet 4.6

$3.00

$15.00

1M

API paid (superseded by Sonnet 5)

Gemini 3.1 Pro

$2.00 (≤200K) / $4.00 (>200K)

$12.00 (≤200K) / $18.00 (>200K)

1M

Limited Gemini app; API paid

Gemini 3.6 Flash

$1.50

$7.50

1M

Gemini app/AI Studio; free API tier + paid API

Gemini 3.5 Flash-Lite

$0.30

$2.50

1M

AI Studio; free API tier + paid API

Qwen 3.7 Max

$1.25 promo / $2.50 list

$3.75 promo / $7.50 list

1M

200 free requests/day; API paid beyond that

MiniMax M3

$0.30 (50% off $0.60)

$1.20 (≤512K)

1M

Open weights; hosting costs apply

LongCat-2.0

Provider-dependent

Provider-dependent

1M

Open weights (MIT); hosting costs apply

NVIDIA Nemotron 3 Ultra

Provider-dependent

Provider-dependent

1M

Open weights (OpenMDW); hosting costs apply

Qwen 3.5 (open-weight)

Self-host / Together

Self-host / Together

1M

Open weights; hosting costs apply

Nex-N2-Pro

Self-host / providers

Self-host / providers

1M

Open weights (Apache 2.0); hosting costs apply

Rio 3.5 Open 397B

Self-host / providers

Self-host / providers

1M

Open weights (MIT); hosting costs apply

Grok 4.3

$1.25

$2.50

1M

Free consumer plan; API paid

Grok 4.5

$2.00

$6.00

500K

Grok Build / Cursor / xAI console; EU partial, API console still closed

Muse Spark 1.1

$1.25

$4.25

1M

Meta Model API ($20 free credits, US preview); free in Meta AI Thinking mode

Kimi K3

$3.00 ($0.30 cache-hit)

$15.00

1M

Free basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License)

Gemini Omni Flash (video)

$1.50

$17.50 (video output)

10-second clips

Gemini app / Flow; AI Studio + API

DeepSeek V4-Pro

$0.435 ($0.0036 cache-hit)

$0.87

1M

DeepSeek Chat free; API paid

DeepSeek V4-Flash 0731

$0.14

$0.28

1M

DeepSeek Chat free; API paid

Kimi K2.7 Code

Provider-dependent

Provider-dependent

256K

Open weights; hosting costs apply

GLM-5.2

Provider-dependent

Provider-dependent

1M

Open weights; hosting costs apply

ERNIE 5.1

China-region pricing

China-region pricing

256K

Baidu free tier

Gemini Spark (agent)

Not API-priced

Not API-priced

1M (Gemini base)

Google AI Ultra $99.99 or $199.99/mo

Fello AI (aggregator)

Routed via app

Routed via app

Model-dependent

$9.99/mo

The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices. OpenAI labels those rows “(<272K context length)” and bills longer prompts at a higher tier, but it no longer publishes the specific long-context figures, so we have stopped quoting them. The GPT-5.6 tiers are billed the same way, at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80.

If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad, starting at $9.99/month with a free tier available. Models are updated regularly so you always have access to the latest.

Best AI for Students & Studying

Best AI for Students & Studying: GPT-5.5 Free + Gemini 3.6 Flash Free (zero-cost frontier for coursework)

The best AI for students is GPT-5.5 Free inside ChatGPT for general coursework and Gemini 3.6 Flash Free inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Opus 5 as the alternative for essay editing.

Most students don’t need to pay: the free ChatGPT tier now defaults to GPT-5.6 (with GPT-5.5 still available), Gemini 3.6 Flash is in the free Gemini app and AI Studio, Claude Sonnet 5 is the new free Claude default, and DeepSeek V4 is free on DeepSeek’s chat site. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI’s new flagship (its FrontierMath score is not yet published, with GPT-5.5 Pro’s verified 39.6% on FrontierMath Tier 4 the current mark), though both are paid-only; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).

Task

Best Model

Why

Free?

Alternative

Essays & coursework

GPT-5.5

Free in ChatGPT, improved factual reliability vs 5.4

Yes

Claude Sonnet 5 (free Claude)

STEM problem-solving

GPT-5.6 Sol / Qwen 3.7 Max

New STEM flagship (5.5 Pro: 39.6% FrontierMath) / 97.1 HMMT 2026 Feb

Pro paid / Qwen API paid

Gemini 3.6 Flash (free)

Research & accuracy

Gemini 3.1 Pro

98% ARC-AGI-1 at $0.52/task, native Google Search grounding

Yes (Gemini app)

Claude Opus 5

Writing editing

Claude Sonnet 5

Free and default on claude.ai; Claude Fable 5 is the quality leader

Yes (Claude free)

Claude Fable 5

Multimodal study (PDFs, slides, images)

Gemini 3.6 Flash

1M context, free in Gemini app

Yes

NotebookLM (Google)

Runner-up and alternatives: Claude Sonnet 5 (free) is the runner-up for essay writing and editing. Gemini 3.6 Flash (free) is the runner-up for multimodal study and PDF ingestion. DeepSeek V4 is the runner-up for problem-solving on a strict zero-cost budget.

Best AI for Work & Professionals

Best AI for Work: GPT-5.6 + Claude Opus 5 ($20/month each, plus Gemini Spark for agents)

The best AI for professional work is GPT-5.6 (ChatGPT’s default since July 9) for daily knowledge work, Claude Opus 5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark on Google AI Ultra at $99.99/month is the only true 24/7 cloud agent.

Use Case

Best Model

Key Stat

Price

Alternative

Daily knowledge work

GPT-5.6

ChatGPT’s default since July 9; the assistant most people can open

$20/mo ChatGPT Plus

Claude Opus 5

Coding (proprietary)

Claude Opus 5

#1 on Arena WebDev (1,702.9) and image-to-WebDev (1,668.6); Anthropic’s recommended default

$20/mo Claude Pro

Claude Fable 5

Coding (cost-effective)

Qwen 3.7 Max

80.4 SWE-Verified, 1M context

$1.25 / $3.75 promo; $2.50 / $7.50 list

DeepSeek V4-Flash 0731

Research & briefings

Gemini 3.1 Pro

98% ARC-AGI-1 at $0.52/task, Google grounding

Google AI Pro / Ultra

Claude Opus 5

Hard math, physics, finance modelling

GPT-5.6 Sol

OpenAI’s new STEM flagship (5.5 Pro verified at 39.6% FrontierMath)

$100/mo ChatGPT Pro

Qwen 3.7 Max

Always-on agent workflows

Gemini Spark

First 24/7 cloud agent

$99.99/mo Google AI Ultra

Claude Cowork

Live news, X-context creative

Grok 4.5

Opus-class + native X grounding

$30/mo SuperGrok

Gemini 3.1 Pro

All-in-one consolidation

Fello AI

ChatGPT + Claude + Gemini + Grok + DeepSeek

$9.99/mo

Pay each vendor separately

Runner-up and alternatives: for most professional teams, Claude Opus 5 is the runner-up to GPT-5.6 for daily work and now the coding leader outright at $5 / $25, with Claude Fable 5 the runner-up for the hardest long-horizon work. Gemini 3.1 Pro is the runner-up for research-heavy roles, and Gemini Spark is the unique pick if you can put a cloud agent to work on long tasks.

Open-Weight and Free Models

Best Open-Weight Models in August 2026: Kimi K3 leads on capability, GLM-5.2 is the one most teams can run

The best open-weight model in August 2026 is Kimi K3, and it took the lead the moment Moonshot published the weights on July 27, 2026. It holds the highest Intelligence Index of any open model at 57 on Artificial Analysis v4.1, six points clear of GLM-5.2, and on Arena it is still the highest-placed open entry, at #2 on WebDev (1,675.5) behind only Claude Opus 5 and #3 on the Agent board.

One catch decides which of the two you should actually use. The K3 download is 96 safetensors shards and about 1.56 TB, which needs a multi-node GPU cluster rather than a workstation, and it ships under a custom Kimi K3 License rather than MIT. So GLM-5.2 (Z.ai, MIT) stays our practical recommendation for teams running their own weights, at Intelligence Index 51, #6 on Arena WebDev (1,587.1) and #5 on the Agent board, with a 1M-token context and a permissive licence.

 

Third place changed hands on the last day of July. DeepSeek V4-Flash 0731 was re-post-trained and jumped from Intelligence Index 40 to 50, which puts it above every model below, at $0.14 / $0.28 under MIT. Behind it the field is crowded and close: DeepSeek V4-Pro and MiniMax M3 tie at 44, Kimi K2.7 Code and Xiaomi’s MiMo-V2.5-Pro at 42, and three models sit level at 41, Nex-N2-Pro plus Hy3 (Tencent, Apache 2.0, July 6) and Inkling (Thinking Machines, July 15). Inkling Small follows at 40 on roughly a quarter of Inkling’s parameters, NVIDIA’s Nemotron 3 Ultra scores 38 and Meituan’s LongCat-2.0 scores 33. Those are v4.1 figures and they are not comparable with the higher numbers Artificial Analysis published earlier in 2026.

 

What changed this month is that the biggest open model stopped being a promise. Moonshot published the full K3 checkpoint on Hugging Face on July 27, 2026, eleven days after the model went live in the Kimi app, at 2.8 trillion total parameters with 104 billion active per token. The licence is worth reading before you build on it, because it is permissive for ordinary use but requires a separate agreement with Moonshot once a model-as-a-service business passes $20 million in revenue over any 12 months. Read our review of Kimi K3.

 

One honest limit on how far open weights have come, and the line moved back this month. Kimi K3 is inside Arena’s top 20 on text overall at #12 and holds #2 on WebDev and #3 on Agent, which is the strongest an open model has placed. But no open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, and it was the last one an open model led. On text, vision, search, document, both coding boards, both image boards and all three video boards, the top entry is proprietary. Open weights are excellent value and genuinely competitive at agentic work; they are not yet at parity across the board.

 

Licensing matters as much as raw score here. Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use, while MiniMax M3 ships under its own community license and Kimi K3 sits on its own custom licence with a revenue threshold for anyone reselling it as a service. Nex-N2-Pro posts the strongest first-party open coding scores at 80.8 on SWE-Bench Verified and 75.3 on Terminal-Bench 2.1.

Model

Best For

Key Benchmark

Context / License

Where To Run

Kimi K3

Highest-scoring open model, agentic and web work

II 57 (v4.1), highest of any open model; #2 Arena WebDev, #3 Arena Agent; 2.8T/104B active

1M / Kimi K3 License

Hugging Face (96 shards, 1.56 TB), Moonshot API, providers

GLM-5.2

Long-horizon agentic coding, 1M context

II 51 (v4.1), highest of any open model you can host on modest hardware; 744B/40B active

1M / MIT

Z.ai, Hugging Face, OpenRouter

DeepSeek V4-Flash 0731

Best value of any model on the page

II 50 (v4.1), up from 40 on July 31; $0.03 per index task, 284B/13B active

1M / MIT

DeepSeek API ($0.14/$0.28), local

LongCat-2.0

Frontier open coder trained on Chinese chips

II 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active

1M / MIT

Hugging Face, GitHub, OpenRouter

MiniMax M3

Cheap frontier-class multimodal

II 44 (v4.1), 59% SWE-Bench Pro, multimodal

1M / license TBD

Hugging Face, API $0.30/1M (50% off)

Nex-N2-Pro

Strongest open coding score

II 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B active

Qwen-based / Apache 2.0

Hugging Face, providers, self-host

Kimi K2.7 Code

Strongest commercially-licensed open coder

+21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active

256K / Modified MIT

Hugging Face, DeepInfra, providers

DeepSeek V4-Pro

Agentic real-world work

II 44 (v4.1), 1.6T/49B active

1M / MIT

DeepSeek API ($0.435/$0.87), local

Hy3

Newest permissive-licence entrant

II 41 (v4.1); #22 on Arena WebDev (1,516.4)

Apache 2.0

Hugging Face, providers, self-host

Inkling

Thinking Machines’ first open model

II 41 (v4.1), agentic 32.3, released July 15

Open weights

Hugging Face, providers, self-host

Inkling Small

Same family at a quarter the size

II 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio in

Apache 2.0

Hugging Face (BF16 and NVFP4), providers, self-host

NVIDIA Nemotron 3 Ultra

NVIDIA-tuned, fully permissive license

II 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active

1M / OpenMDW

OpenRouter, Hugging Face, AWS (8× B200 self-host)

Qwen 3.5 (397B / 17B active)

Multimodal, fast decode

88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v6

1M / open

Together, OpenRouter, local

Qwen3.6-35B-A3B

Efficient open agentic coder (3B active)

86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active

262K (→1M YaRN) / Apache 2.0

Hugging Face, OpenRouter, local

Qwen3.6-27B

Laptop-runnable dense coder

87.8 GPQA Diamond, dense 27B, multimodal

256K / Apache 2.0

Local Mac/PC, Hugging Face, OpenRouter

Rio 3.5 Open 397B

Qwen 3.5 fine-tune, multilingual reasoning

70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5

397B / 17B active, MIT

Hugging Face, providers, self-host

Qwen 3.5-9B

Laptop-runnable open-weight

81.7 GPQA Diamond

Dense / open

Local Mac/PC with 16GB+ RAM

Llama 4 Maverick

Meta-line flagship

17B active / 400B total params

Llama 4 license

Meta cloud, Hugging Face, local

NVIDIA Nemotron 3 Nano Omni

Edge / low-power

Multimodal, very small footprint

Compact / open

Local, NVIDIA tool

Runner-up and alternatives: GLM-5.2 is the runner-up on score and the one to pick if you are hosting the weights yourself, DeepSeek V4-Flash 0731 is third at 50 and the cheapest way to get a 1M-context open model by a wide margin, DeepSeek V4-Pro and MiniMax M3 follow at 44, Hy3 and Inkling sit at 41 with Inkling Small just behind at 40, and NVIDIA Nemotron 3 Nano Omni is the natural pick for edge and on-device use.

How We Evaluate

Benchmarks, Prices, and Hands-On Use

Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 2.1 board, covering the Intelligence and Agentic indexes, GPQA Diamond, ARC-AGI-1 through 3, Humanity’s Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor’s official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.

Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis’s professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We do not move a category crown on the strength of a single leaderboard, and where a new model is too recent for the human-preference boards to have rated it, we say so rather than crowning it on benchmark scores alone. When a model goes through a major upgrade between updates, we re-rank the category and add a “What changed this month” line at the bottom of the deep-dive.

FAQ

What is the best AI model right now in August 2026?

It depends on the task. On overall benchmark score, Claude Opus 5 (July 24) is #1 on Artificial Analysis’s Intelligence Index at 61 and its Agentic Index at 55.3, and Arena has now rated it top of four of its boards, including both coding boards. For daily chat, GPT-5.6 has been ChatGPT’s default since July 9 and is the assistant most people can actually open, although Claude Fable 5 leads Arena’s text leaderboard. For coding, Claude Opus 5 is #1 on Arena’s WebDev (1,702.9) and image-to-WebDev (1,668.6) boards at half Fable 5’s price. For writing, Claude Fable 5 is #1 on Arena’s text leaderboard (1508.6), Humanity’s Last Exam (53.3%) and AA-Omniscience (40). For accuracy, Gemini 3.1 Pro ties the human panel on ARC-AGI-1 at 98% for $0.52 a task. For hard math and reasoning, GPT-5.6 Sol is #1 on LiveBench Mathematics, LiveBench Reasoning and ARC-AGI-2. For images, ChatGPT Images 2.0 leads both boards, and for video Gemini Omni Flash leads both. For agents, Gemini Spark is the 24/7 cloud agent and Claude Cowork the desktop one.

What is new in AI in August 2026?

August opens on two crown changes rather than a new flagship. Claude Opus 5 took the coding crown after Arena finished rating it #1 on both WebDev (1,702.9) and image-to-WebDev (1,668.6), and DeepSeek V4-Flash 0731 (July 31) took price-performance by reaching Intelligence Index 50 at $0.14 / $0.28 per 1M tokens. MiniMax H3 (July 31) added 2K video with native stereo audio from $0.13 a second. Looking back at July, the biggest launch was Claude Opus 5 (July 24), which took #1 on Artificial Analysis’s Intelligence Index at 61 and Agentic Index at 55.3 at $5 / $25 per 1M tokens.

Moonshot launched Kimi K3 on July 16 at Intelligence Index 57 and published its open weights on July 27, and Google shipped Gemini 3.6 Flash on July 21, cutting output to $7.50 per 1M tokens. The month opened with Claude Fable 5 back online (July 1) after the US government lifted the June 12 export-control order that had pulled it offline. On July 8 and 9, four more launches landed: OpenAI’s GPT-5.6 family (Sol, Terra, Luna) reached general availability on July 9 and became ChatGPT’s new default; xAI took Grok 4.5 public on July 8 as a cheap Cursor-trained coding model at $2 / $6 per 1M tokens that Artificial Analysis scores at Intelligence Index 54 on v4.1; Meta shipped Muse Spark 1.1 on July 9 as its first paid model, a $1.25 / $4.25 agentic coder; and ByteDance released Seedream 5.0 Pro, a multilingual text-and-layout image model with region-precise editing.

The rest of the headlines landed in the final week of June: OpenAI first previewed GPT-5.6 on June 26 behind a US-government access list; Meituan open-sourced LongCat-2.0, a 1.6-trillion-parameter MIT coding model trained entirely on Chinese chips, on June 29; and Anthropic made Claude Sonnet 5 its new default model on June 30. Gemini 3.5 Pro is still unreleased; Bloomberg reported on July 16 that it is months behind schedule and short of Google’s internal goals, particularly in coding. These follow the late-May and June board of Claude Opus 4.8 (Intelligence Index 56 on v4.1), Qwen 3.7 Max, Gemini 3.5 Flash, Gemini Spark, and the open-weight wave of MiniMax M3, NVIDIA Nemotron 3 Ultra, Kimi K2.7, and GLM-5.2.

What is Claude Opus 5?

Claude Opus 5 is Anthropic’s newest flagship, released July 24, 2026, and the current #1 model on Artificial Analysis. It leads both the Intelligence Index at 61 and the Agentic Index at 55.3, and tops the GDPval-AA v2 professional-deliverables board at 1858, well clear of Claude Fable 5’s 1746. API pricing is $5 / $25 per 1M tokens, the same as Opus 4.8 and half the cost of Fable 5, with a 1M-token context window and a May 2026 knowledge cutoff. ARC Prize independently confirms Anthropic’s launch claim on ARC-AGI-3, where Opus 5 scores 30% against 8% for the next-best model, roughly 3.75x. Arena has since rated it, placing it #1 on WebDev, image-to-WebDev, Document and the Agent board and #6 on text, and that is what moved the coding crown to it. One thing to keep in mind: Artificial Analysis measures its hallucination rate at 50%, which is why it holds no accuracy crown on this page. Read our cover of Claude Opus 5.

Is Claude Fable 5 back?

Yes. Anthropic redeployed Claude Fable 5 on July 1, 2026 after the US government lifted the export-control restriction it had imposed on June 12. It is available again on the Claude API, Claude.ai, Claude Code, and Claude Cowork. Anthropic made Fable 5 permanent in the paid plans on July 20, 2026. Max and Team Premium include it at roughly 50% of regular usage limits; Pro and Team Standard reach it through usage credits with a one-time $100 starting credit. API pricing is $10 / $50 per million tokens. Fable 5 is a Mythos-class model built for long-horizon agentic work with a 1M-token context, and it is the runner-up for coding behind Claude Opus 5, at #2 on Arena’s image-to-WebDev board (1,625.7) and #4 on WebDev.

What is Claude Sonnet 5?

Claude Sonnet 5 is Anthropic’s new default model, launched June 30, 2026 for Free and Pro users on claude.ai and live in the Claude API, Claude Code, Cursor, VS Code, and GitHub Copilot. It ships with a 1-million-token context window at introductory pricing of $2 / $10 per 1M tokens through August 31, 2026 (then $3 / $15). It scores 1,603 on GDPval-AA v2, ahead of Opus 4.8’s 1,593, and closes much of the agentic gap to Opus 4.8 while costing far less than GPT-5.5. Note that those Elo boards are re-fitted as models are added, so these numbers are lower than the ones we published earlier and Claude Opus 5 now tops the board at 1,861. Sonnet 5 no longer holds our writing crown, which moved to Claude Fable 5 once we added the human-preference boards, but it is still the writing model most people should use because it is free and default on claude.ai. The one caveat is an updated tokenizer that maps the same text to roughly 1.0-1.35x more tokens.

What is GPT-5.6 and can I use it?

Yes, as of July 9, 2026. GPT-5.6 is OpenAI’s next-generation model family, and after a two-week gated preview that began June 26 behind a US-government safety review, it reached general availability on July 9. There are three tiers, least to most capable: Luna, the fast, cheapest tier ($0.20 / $1.20 per 1M tokens); Terra, a balanced everyday model OpenAI says matches GPT-5.5 ($2 / $12); and Sol, the flagship tuned for biology, chemistry, and cybersecurity ($5 / $30). OpenAI cut Luna by 80% and Terra by 20% on July 30, 2026, and left Sol and every ChatGPT subscription price untouched. It is now live across ChatGPT, Codex, and the API as OpenAI’s default model. On the limited benchmarks OpenAI published, Sol scores 88.8% on Terminal-Bench 2.1 (91.9% in ultra mode) versus GPT-5.5’s 88.0%; OpenAI withheld the usual SWE-bench, GPQA, and FrontierMath numbers, and the context window is still not officially published. One caveat: OpenAI’s system card and the external evaluator METR flagged elevated “scheming” behaviour in Sol, so treat it carefully for high-stakes factual work until independent results settle.

Is Grok 4.5 out yet?

Yes. xAI released Grok 4.5 publicly on July 8, 2026, its first flagship since SpaceX absorbed the company and went public as SPCX. Elon Musk describes it as “an Opus-class model, but faster, more token-efficient and lower cost,” and says internally it is “roughly comparable to Opus 4.7, but much faster.” It is Cursor-trained and aimed at coding and agentic work, priced at $2 / $6 per 1M tokens with a 500K-token context window, well under Opus 4.8’s $5 / $25. It is live in Grok Build, Cursor on all plans, and the xAI console. EU access began rolling out after a July 16 announcement and is still partial, with Cursor reporting availability while xAI’s API console remains closed to EU users under the AI Act’s systemic-risk obligations. Independent scores are now in: Artificial Analysis scores it at Intelligence Index 54 on v4.1, matching the field on Terminal-Bench 2.1 (83.3%) but trailing on SWE-Bench Pro at 64.7% (below Opus 4.8’s 69.2%), so it is the value pick rather than the outright benchmark leader. Grok 4.5 is also our creativity pick, but on product grounds rather than quality: it has the fewest content restrictions and native real-time X access, while sitting #33 on EQ-Bench Creative Writing and #41 on Arena’s creative-writing board. For the best-written output, Claude Fable 5 wins that comparison outright.

What is the best open-weight AI model in 2026?

Kimi K3 is the best open-weight model right now, and it has been since Moonshot published the weights on July 27, 2026. It has the highest Intelligence Index of any open model at 57 on Artificial Analysis v4.1, and on Arena it is #2 on WebDev and #3 on the Agent board. The catch is that it is 96 shards and about 1.56 TB under a custom Kimi K3 License, so GLM-5.2 (June 13, MIT) remains the model most teams can actually host, at Intelligence Index 51 and #6 on Arena WebDev. Third is DeepSeek V4-Flash 0731, which jumped from Intelligence Index 40 to 50 on July 31 without changing price, size or licence, making it the best value in the open field at $0.14 / $0.28 under MIT. Behind them, DeepSeek V4-Pro and MiniMax M3 tie at 44, Kimi K2.7 Code and MiMo-V2.5-Pro at 42, and three sit at 41, Nex-N2-Pro plus Hy3 (Tencent, Apache 2.0) and Inkling (Thinking Machines), with Inkling Small at 40. Nemotron 3 Ultra scores 38 and LongCat-2.0 scores 33. Worth knowing honestly, no open model is first on any Arena board we track any more.

What is Qwen 3.7 Max and how does it compare to GPT-5.5?

Qwen 3.7 Max is Alibaba’s flagship API model, launched May 20, 2026 at the Alibaba Cloud Summit in Hangzhou. It scores Intelligence Index 46 on Artificial Analysis’s current v4.1 scale (it scored 57 at launch on the older scale), 92.4 on GPQA Diamond, 97.1 on HMMT 2026 February, and 80.4 on SWE-Verified. List API pricing is $2.50 / $7.50 per 1M tokens with a 1M-token context window, plus $0.25 cached input (a 90% cache discount); a 50% launch promotion is still active, cutting it to $1.25 / $3.75 (cached input $0.125), and Alibaba now includes 200 free model requests per day. Compared to GPT-5.5 at $5 / $30, Qwen 3.7 Max is half the input cost and a quarter of the output cost, but GPT-5.5 still leads on Intelligence Index (55 versus 46 on v4.1) and on FrontierMath. For cost-sensitive agentic and long-context work where you want frontier-adjacent quality, Qwen 3.7 Max is the value pick.

What is GPT-5.5 and how is it different from GPT-5.4?

The GPT-5.5 family launched April 23, 2026. The headline change is factual reliability: on a selected set of user-flagged conversations, OpenAI reports that GPT-5.5’s individual claims were 23% more likely to be factually correct than GPT-5.4’s, with full responses containing a factual error about 3% less often, plus faster response times across all tiers and a refreshed memory system. The consumer default is GPT-5.5 Instant (free with limits), and the gpt-5.5 API model runs $5 / $30 per 1M tokens. GPT-5.5 Pro is the higher-reasoning variant, available inside ChatGPT Pro at $100/month, and leads FrontierMath Tier 4 at 39.6%. Its successor GPT-5.6 reached general availability on July 9, 2026. OpenAI still sells GPT-5.5 at $5 / $30, but Artificial Analysis has since marked every GPT-5.5 reasoning tier deprecated, so treat it as a model you can still buy rather than a current benchmark reference.

Is ChatGPT still the best AI?

Not on every benchmark, but it is still the best default. GPT-5.6 became ChatGPT’s default on July 9 and is the assistant most people can open, though Claude Fable 5 leads Arena’s text leaderboard on measured preference. GPT-5.5 Instant is still the safer pick for hallucination-sensitive work. Claude Opus 5 is the better pick for coding and long agentic tasks and leads Intelligence Index v4.1 at 61, ahead of Claude Fable 5 at 60. Gemini 3.1 Pro is the better pick for accuracy and research. Gemini 3.6 Flash is the better pick for price-performance. Qwen 3.7 Max is the better pick for cost-effective frontier work. ChatGPT remains the most polished consumer product overall and the natural starting point if you only pay for one model.

What is Gemini Spark and is it worth $99.99/month?

Gemini Spark is Google’s first 24/7 cloud-resident AI agent, launched at Google I/O on May 19, 2026 and exclusive to the Google AI Ultra plan, which Google restructured at I/O to include a $99.99/month entry tier and a $199.99/month top tier. Spark is built on Gemini base models with Google’s Antigravity harness on a Google Cloud VM, integrates with Gmail, Google Docs, and other Google Workspace apps, and can interact with Chrome and Android’s Halo system on the device side. It is worth the spend for users who have repeatable long-running workflows (inbox triage, research roll-ups, scheduled tasks). For one-off tasks, Claude Cowork at $20/month covers most desktop-agent needs.

What is the cheapest frontier-class AI model?

On API pricing per million tokens, GPT-5.6 Luna is now the cheapest closed frontier-class model at $0.20 / $1.20, after OpenAI cut it 80% on July 30, 2026, and Artificial Analysis scores it at Intelligence Index 51, a point above Gemini 3.6 Flash. Cheaper still is DeepSeek V4-Flash 0731 at $0.14 / $0.28, which is now our price-performance pick: it matches Gemini 3.6 Flash’s Intelligence Index of 50 and costs $0.03 per index task against Flash’s $0.56, under an MIT licence you can self-host. The caveat is that the score is one day old, comes from a single house and has no Arena votes behind it. MiniMax M3 now lists at $0.30 per million input tokens on a permanent 50% discount, with LongCat-2.0 as a strong MIT alternative. Qwen 3.7 Max at $1.25 / $3.75 on its current promo ($2.50 / $7.50 list) is the value pick at Intelligence Index 46, though Luna now undercuts it on both price and score, and for an open-weight model with a 1M context, DeepSeek V4-Flash at $0.14 / $0.28 is the cheapest by an order of magnitude. Note that OpenAI bills prompts over 272K input tokens at a higher rate. It no longer publishes the GPT-5.5 and GPT-5.5 Pro long-context figures, but the GPT-5.6 tiers are charged at 2x input and 1.5x output for the whole request, so a long-context Terra call runs $4 / $18 and Luna $0.40 / $1.80.

Which AI models are free?

ChatGPT Free now defaults to GPT-5.6 (with GPT-5.5 still available) under usage limits. Gemini Free runs Gemini 3.6 Flash in the Gemini app and Google AI Studio. Claude Free runs Claude Sonnet 5 (the new default) with daily limits. DeepSeek Chat runs DeepSeek V4 free on the DeepSeek website. Grok has a limited free consumer plan (X Premium is a paid add-on). Qwen 3.5, NVIDIA Nemotron 3 Ultra, MiniMax M3, LongCat-2.0, Kimi K2.6, DeepSeek V4, and GLM-5.2 are open-weight and free to self-host. Qwen 3.7 Max is API-only with no consumer chat front-end, but Alibaba now includes 200 free model requests per day. Kimi has a free basic tier in its app, with heavier agentic use metered.

Which AI is the best for coding?

Claude Opus 5 is the best for coding, and it wins the two boards where developers vote on the result rather than a harness running a script: #1 on Arena’s WebDev board (1,702.9) and #1 on image-to-WebDev (1,668.6), on August 1 and July 31 vote cutoffs. It runs $5 / $25 per 1M tokens, half of Claude Fable 5, and Anthropic’s own docs say to start with Opus 5 for complex agentic coding. Claude Fable 5 is the runner-up for the hardest long-horizon work at #2 image-to-WebDev and #4 WebDev, and Kimi K3 is the contender at #2 on WebDev (1,675.5). Note that Artificial Analysis’s Coding Index is not a Claude sweep: GPT-5.6 Sol (xhigh) leads it at 78.3 with Opus 5 at 78.0, close enough to call a tie. DeepSeek V4-Flash 0731 is the price-performance pick at $0.14 / $0.28, and GLM-5.2 (MIT) is the best open-weight coder you can host yourself, with Kimi K3 scoring higher but needing a multi-node cluster.

Which AI is the best for writing?

Claude Fable 5 is the best for writing, and it is the only model in the top three of all three independent writing boards: #1 on Arena creative writing, #1 on LiveBench Language (90.7) and #3 on EQ-Bench Creative Writing v3. It costs $10 / $50 per 1M tokens. Claude Sonnet 5 is the value pick and what most people should actually use, since it is free and default on claude.ai at $2 / $10 introductory pricing, though it ranks #53 on Arena creative writing. Kimi K3 wins EQ-Bench outright at 2377 if you want distinctive fiction, Claude Opus 5 tops the GDPval-AA v2 professional-deliverables board at 1858, GPT-5.5 is the alternative for fact-anchored business writing, and Gemini 3.6 Flash is the price-performance pick for bulk content.

Which AI is the best for accuracy and research?

Gemini 3.1 Pro is the best for accuracy and research. It ties the human panel on ARC-AGI-1 at 98% and does it at $0.52 per task, scores 94.1% on GPQA Diamond where GPT-5.6 Sol now matches it, and has native Google Search grounding for live factual answers. Its 77.1% on ARC-AGI-2 is still correct but now ranks around 14th, so we no longer lead with it. Two caveats worth knowing: on grounded search specifically, Arena’s leaderboard is led by Anthropic rather than Google, with Gemini 3.1 Pro grounding at #7; and on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 at 93% while Claude Opus 5 leads ARC-AGI-3 at 30%. We did not move this crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50%. Gemini 3.5 Pro is still unreleased.

Which AI is the best for images?

ChatGPT Images 2.0 is the best for images, and it leads both boards: #1 on Arena text-to-image at 1385 and #1 on image editing at 1463. It is still the strongest at readable multilingual text and infographic-style output. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, with Reve 2.0 now behind it on both boards. Meta’s Muse Image is #3 on Arena’s text-to-image board and #2 on image editing, though on Artificial Analysis’s image board third place goes to Microsoft’s MAI-Image-2.5 at 1269.7. Google’s Nano Banana Pro is not the runner-up: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, which is the photoreal pick. ByteDance’s new Seedream 5.0 Pro (July 8) is a strong multilingual-text and region-precise-editing option, though it has no independent benchmarks yet and is accessed via BytePlus and Magnific rather than a consumer app. Midjourney v8 is the best for stylized art, and Grok Imagine is the only frontier model that allows Spicy Mode adult content.

Which AI is the best for video?

Gemini Omni Flash is the best for AI video and it is #1 on both independent video boards, leading Arena’s text-to-video leaderboard at 1527 Elo, 45 points clear of second place, and Artificial Analysis’s video arena at 1246.6. It costs roughly $0.10 per second of finished video and generates 10-second clips with conversational editing. ByteDance’s Dreamina Seedance 2.0 is the runner-up at 1482 and actually beats it on image-to-video, and Meta’s Muse Video is third. Veo 3.1 is the pick when you need longer clips, but it is not the quality leader, sitting #6 on Arena’s video board. OpenAI retired the Sora 2 consumer app on April 26, 2026 and its API runs until September 24, 2026. Read our breakdown of Gemini Omni Flash.

Which AI is the best for hard math and STEM problems?

GPT-5.6 Sol is the best for hard math and STEM, and this is the best-supported pick on the page: #1 on LiveBench Mathematics (96.2), #1 on LiveBench Reasoning (91.7) and #1 on ARC-AGI-2 (93%) against a 100% human panel. It is included in ChatGPT Pro at $100/month. OpenAI has still not published Sol’s FrontierMath score, so GPT-5.5 Pro’s verified 39.6% on Tier 4 remains the cited OpenAI mark. Qwen 3.7 Max leads competition math at 97.1 on HMMT 2026 February and 44.5 on Apex at a fraction of the cost, with 200 free requests a day. Claude Opus 5 is the alternative for long agentic reasoning chains, leading the Agentic Index at 55.3 and ARC-AGI-3 at 30%.

Which AI is the best for creativity?

Grok 4.5 is our creativity pick, and we want to be exact about why: it is the product call, not the quality call. It has the fewest content restrictions of any frontier model and the only native real-time X integration, at $30/month on SuperGrok. It is not the better writer, sitting #33 on EQ-Bench Creative Writing and #41 on Arena’s creative-writing board, losing both to the older Grok 4.20-beta1. If you are choosing on output quality, Claude Fable 5 is #1 on Arena creative writing and Kimi K3 wins EQ-Bench outright. Claude Opus 5 is the alternative for structured long-form creative work, and Gemini 3.1 Pro for multimodal creative.

What is Fello AI?

Fello AI is an AI chatbot for Mac, iPhone, and iPad that lets you use all top AI models like ChatGPT, Claude, Gemini, Grok, and DeepSeek in one app, with models updated regularly so you always have the latest. It is $9.99/month with a 4.7-star rating across 27,000+ reviews.

How often do you update this page?

We update this page at least monthly and within 24-48 hours of any major model launch. 

Fello AI macOS app interface showing an AI chat workspace with file attachments, image generation, document analysis, and bookmarked conversations in a dark desktop UI.

Download Fello AI,
the all-in-one AI App

Use all the latest AI models like ChatGPT, Gemini, Claude or Grok in one app!

rating 4.7, 27K+ reviews