Updated September 19, 2026

Best AI Models in 2026

Rankings, comparisons, and deep dives, updated monthly as new models ship.

OpenAI launched GPT-6 Astra on September 3, the biggest launch of the month and, by six tenths of a point, still not the biggest score. Artificial Analysis now measures it at 52.8 on Intelligence Index v4.3, 5.7 points clear of GPT-5.6 Sol and 0.6 behind Claude Fable 5.1, at $10 / $50 per 1M tokens against Sol's $4 / $20. Where it does move is coding: it leads Artificial Analysis's own run of Terminal-Bench 4.0 at 60% against Claude Fable 5.1's 55%, both at xhigh effort, and it gets there on roughly 17,000 output tokens a task against Fable 5.1's 55,000. It is the first model OpenAI has designated Critical for cybersecurity, so approved defenders in its Daybreak program went first; ChatGPT Business and Pro followed on September 4 and Plus a few hours later, leaving the API, AWS, Azure and the free tier still to come. This page ranks on measured scores, so a model takes a crown as soon as it out-scores the holder on the boards that define the category, without waiting for the vote-based boards to catch up. A model nobody can use yet is flagged on the card rather than denied one.

Claude Fable 5.1 arrived on September 1 and went straight to the top of Artificial Analysis's Intelligence Index, where it still leads at 53.4 on v4.3, and it tops AA-Briefcase at 1646 as well. Claude Opus 5 becomes the value pick on the thing most readers are actually choosing on, the price, $5 / $25 per 1M tokens against Fable 5.1's $10 / $50. The vote-based boards have since gone to OpenAI: GPT-6 Astra leads Arena's Code Arena: WebDev at 1800 and image-to-WebDev at 1733, with Fable 5.1 second on both and Claude Opus 5 third, while Alibaba's Qwen3.8-Max-0902, which debuted at #1 on WebDev on September 2, has slipped to fourth. Fable 5.1 now has thousands of Arena votes, so the human-preference boards have had their say. Grok 4.6 was August's real surprise: a post-training refresh of Grok 4.5 that gained five points, 44.4 against 39.1, while charging $2 / $6, and finishing long agentic jobs in roughly half the turns Opus 5 needs. Meta is climbing faster than anyone: Muse Spark 1.3 landed on September 2 at 48.2 on the same index, behind Claude Opus 5 and Claude Fable 5, though it shipped to developers first, on Muse Code and the Meta Model API, with Meta saying a Facebook, Instagram and Meta AI rollout would follow. Our guide to AI benchmarks explains how that index is built and why its version number matters.

The other move is at the cheap end, and this month it went both ways. DeepSeek raised both V4 models roughly fourfold on August 16 and split the rate card into peak and off-peak tiers, which cost DeepSeek V4-Flash 0731 the price-performance pick; on September 10 it reversed most of that, retiring V4-Flash and replacing it with DeepSeek-V4.1-Flash at $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak, MIT weights on the day. The pick stays with GLM 5.3 Flash, which landed on August 26 with MIT weights on day one and now sells at its plain $0.15 / $0.50 list, the launch promotion having expired on September 9: Artificial Analysis measures it at $0.25 an Intelligence Index task for a score of 41.9, against DeepSeek's $0.27 for 39.5. Off-peak the two are level on input and GLM is cheaper on output; at peak GLM wins both. Google had already halved its own Flash rate: Gemini 3.7 Flash shipped on August 13 at an introductory $0.75 / $3.75, with Gemini 3.6 Flash moved onto the same rate, and Gemini 3.8 Flash followed on September 2 at exactly that price, scoring 41.2 on Artificial Analysis's v4.3 index against 39.4 for 3.7 Flash. All three prices double on January 1, 2027.

The open field then closed the month with three releases in three days: GLM 5.3 finally published its weights on August 28, though under a bespoke Z.ai licence rather than the plain MIT its smaller sibling got, Alibaba opened Qwen3.8-Flash-Next on August 26 as the first public look at the Qwen4 architecture, and Tencent open-sourced Hy4 Preview on August 28, 770 billion parameters under Apache 2.0. Below are the category winners for September 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks, and pricing are updated within 48 hours of any major model launch.

September 2026 Category Winners
Writing
Claude Fable 5.1
#1 Intelligence Index (53.4) and #1 Humanity's Last Exam (59%)
Beats Claude Fable 5 on every board that has measured both, at the same $10 / $50. Best value: Claude Sonnet 5, free on claude.ai.
Deep dive →
Chat & Daily Assistant
GPT-5.6
ChatGPT default since July 9
The best assistant most people can actually open, balancing capability, speed, and reach.
Deep dive →
Images
ChatGPT Images 2.5
#1 and #2 on both of Arena's image boards
Its two models took text-to-image and image editing within a day of launch, and it is the default on every ChatGPT tier including free. Best value: Flare, the faster of the two, at the same token price as GPT Image 2.
Deep dive →
Video
Gemini Omni 1.1 Flash
#1 Arena text-to-video (1515), 40-second scenes
Google's newest video model: scene extension to 40 seconds, 4K output, from $0.03 a second.
Deep dive →
Coding
GPT-6 Astra
60% Terminal-Bench 4.0 and #1 on both Arena coding boards
Leads Artificial Analysis's own Terminal-Bench 4.0 and both of Arena's coding boards, at $10 / $50. Best value: Claude Opus 5 at $5 / $25.
Deep dive →
Creativity
Grok 4.6
Fewest content restrictions + native real-time X
The pick for edgy, on-trend work: fewest guardrails and native X integration.
Deep dive →
Accuracy & Research
Claude Fable 5.1
Highest factual accuracy Artificial Analysis has measured (67%)
Tops the boards that measure how often a model is simply wrong. Best value: Gemini 3.1 Pro at $0.52 a task with Google Search grounding.
Deep dive →
Problem Solving
GPT-6 Astra
#1 LiveBench Reasoning (92.65) and #1 ARC-AGI-2 (95%)
Took the hardest reasoning boards off GPT-5.6 Sol days after launching. Best value: GPT-5.6 Sol at $4 / $20, still second on ARC-AGI-2.
Deep dive →
AI Agents
Gemini Spark
First 24/7 cloud-resident AI agent
Runs continuously in a Google Cloud VM, deeply integrated with Gmail, Docs, and Chrome.
Deep dive →

Want the top AI models without juggling separate subscriptions? Fello AI brings the leading models together in one native app for Mac, iPhone and iPad.

Download Fello AI
What's New in September 2026
Sep 15 TypeSafe TypeSafe leaves stealth with Jev, a model that returns typed decisions instead of text, at $0.042 per 1M tokens New
TypeSafe AI came out of two years of stealth on September 15, 2026 with $40 million in seed funding led by DCVC and a model that does not generate text at all. Jev is the first of what the company calls System One models: you send a block of state plus a fixed set of typed questions, and it answers all of them in one parallel pass, returning a choice, a score or a yes-no probability instead of a sentence. Because the answer space is defined before the call it cannot return a malformed result, which is where the launch post's line that Jev can't hallucinate comes from. It can still pick the wrong option, and TypeSafe's own documentation says calibration does not guarantee that an individual answer is correct. Pricing is $42 per billion input tokens, which is $0.042 per 1M, with output free, on a 64k context and text input only. The performance claims need care, because the company published four different ones in three days: from the press release's under 100 milliseconds and up to 100x faster to the homepage's 193.6x faster and 444.6x cheaper, each measured against a different rival, with its own footnote calling the big numbers the higher end of real world gains. Its workflow eval scores against the average of GPT-6 Astra and Fable 5.1 rather than ground truth, and it publishes no public benchmark scores at all by choice. Two independent early-access tests on September 18 measured roughly 6x cheaper and 8x faster against a cheap model with reasoning off, and the calibration held up. Nothing here moves a pick on this page: Jev has no consumer surface and is waitlisted early access. Full detail in our System One models explainer.
Sep 10 DeepSeek DeepSeek retires V4-Flash, replaces it with V4.1-Flash and cuts the Flash rate by about a third New
DeepSeek retired DeepSeek-V4-Flash on September 10, 2026 and put DeepSeek-V4.1-Flash in its place. The model name is now deepseek-flash; the legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names still resolve, but the models behind them have been retired and the requests are served by V4.1-Flash and billed at the Flash price. That price is a cut of roughly a third: $0.15 / $0.60 per 1M tokens off-peak and $0.30 / $1.20 at peak, against $0.22 / $0.66 and $0.44 / $1.32 for the build it replaces, with cache hits at $0.003 off-peak. The peak windows are unchanged, 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, with every other hour at half. V4-Pro is untouched at $0.66 / $1.98 off-peak, and DeepSeek has said it will keep serving it past its September 14 retirement date. The new model is a 552-billion-parameter mixture-of-experts on a causal encoder-decoder design that activates 8 billion parameters at prefill and 16 billion at decode, takes a 1M-token context, reads images natively, and went up on Hugging Face under MIT the same day. Artificial Analysis scores it 39.5 on Intelligence Index v4.3 against 34.5 for the V4-Flash 0731 build it replaces, at $0.27 an index task. It does not retake the price-performance pick: GLM 5.3 Flash scores 41.9 at $0.25 a task. Full detail in our DeepSeek pricing guide.
Sep 3 OpenAI GPT-6 Astra ships at $10 / $50, and within a day it is on Plus and Pro and leading the coding boards New
OpenAI launched GPT-6 Astra on September 3, 2026, settling the argument that ran all year: this is GPT-6, not another point release inside the GPT-5 line. Company president Greg Brockman told a press briefing "Welcome to the AGI era." The independent numbers are more measured than the framing. Artificial Analysis scores Astra 61 on Intelligence Index v4.1.1 at max effort, exactly level with GPT-5.6 Sol, five points behind Claude Fable 5.1 at 66 and two behind Claude Opus 5 at 63, and it charges $10 / $50 per 1M tokens against Sol's $4 / $20, which works out 75% more expensive per index task than the model it replaces. The stronger result is on the Coding Agent Index, where Astra in Codex scores 67, level with Claude Opus 5 and Claude Fable 5 in Claude Code and behind Claude Fable 5.1's 70, and it gets there on one third of Sol's tokens and one fifth of Opus 5's, which puts it on the cost-efficiency frontier. It also halves hallucination on AA-Omniscience, from 92% to 51% at max effort, while gaining four points of accuracy, and adds roughly 80 Elo on AA-Briefcase and six points on Humanity's Last Exam. Against that it drops about 80 Elo on GDPval-AA v2 and regresses two to three points on customer support, SciCode and long-context reasoning. Two caveats travel with the launch numbers: the 98.6% on ARC-AGI-3 is an action-efficiency score against a human baseline rather than a solve rate, measured on a harness OpenAI itself has shown can triple the result, and Epoch AI discloses that OpenAI funded FrontierMath and holds exclusive access to part of it. Availability is the real constraint. Astra is the first model OpenAI has ever designated Critical for cybersecurity under its Preparedness Framework, so approved defenders in the Daybreak program got it first. ChatGPT Business and Pro followed on September 4 and Plus a few hours later, with the API, AWS, Azure and the free tier still outstanding. Artificial Analysis has since retired the Coding Agent Index, and on its own Terminal-Bench 4.0 run Astra now leads outright, 60% to Claude Fable 5.1's 55%, which is what moved the coding crown on this page to Astra. Full detail in our GPT-6 Astra breakdown.
Sep 2 Alibaba Qwen3.8-Max-0902 takes Arena's WebDev board off Claude Opus 5 by three points, at $2 / $6 New
Alibaba shipped Qwen3.8-Max-0902 on September 2, 2026, and the naming is the first thing to get right: this is a post-training refresh of the Qwen3.8-Max that launched on August 3, not a new version, running the same 2.4 trillion parameters and the same 1M-token context window. Qwen says it was further post-trained on coding and Cowork-style work. The reason it matters is the board it moved. Arena announced the same day that Qwen3.8-Max-0902 debuted at #1 overall on Code Arena: WebDev with 1691 points, three points above Claude Opus 5 (Max), seventeen above Kimi K3 (Max) and twenty-two above the previous Qwen3.8-Max, taking #1 in the Data & Analytics and Consumer Product categories as well. That is the first time this year an Anthropic model has been displaced at the top of a vote-based coding board. Pricing is $2 / $6 per 1M tokens with cache hits at $0.17 explicit and $0.25 implicit, which Arena scores as a blended $5 per million and the best position on its Pareto frontier. Read it with three caveats. Three points is inside the noise on a vote-based board, Artificial Analysis had published no Intelligence Index for the 0902 snapshot at the time and has since scored it 45.4 on v4.3, and this is a hosted model on QwenCloud with no consumer chat front-end: the open weights Alibaba promised for Qwen3.8-Max in August still have not shipped, so nothing here changes our open-weight pick. Arena says Agent Arena scores are still to come. Qwen3.8-Max-0902 has since slipped to fourth on WebDev, behind GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5, and the coding category has moved to Astra. Full detail in our Qwen 3.8 breakdown.
Sep 2 Meta Muse Spark 1.3, Intelligence Index 57 to 61 at an unchanged $1.25 / $4.25, winning coding and losing every agent row New
Meta released Muse Spark 1.3 on September 2, 2026, its fourth Muse Spark model in five months, in Muse Code and the Meta Model API. Artificial Analysis scores it 61 on Intelligence Index v4.1.1, up from 57 for 1.2 and 53 for 1.1, which is eight points in two months and the fastest climb anyone on this page is managing. Pricing did not move at all: $1.25 / $4.25 per 1M tokens on a 1M-token context window, with the Contributor tier still at $0.10 / $0.20 in exchange for letting Meta train on your prompts and completions. Two things about the launch need reading carefully. On Meta's own scorecard the model wins every coding row, DeepSWE v1.1 at 75.4 against Claude Opus 5's 74.0 and SWEAtlas CodeBase QnA at 59.4, ties GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8, and takes both MRCR long-context bands by a distance, 98.1 on the 512K to 1M band against 55.5 for 1.2. It also loses all six rows Meta groups under Agent, four to Opus 5 and two to GPT-5.6 Sol, and that half of the chart went almost unreported. The other catch is which model was measured: the benchmarked column is Muse Spark 1.3 (max), a limited preview for Meta's partners that scores 62, while the version anyone can call today is xhigh at 61, and Meta ran the 1.2 comparison column at xhigh rather than max. Meta shipped it to developers first, on Muse Code and the Meta Model API, and said a rollout to Facebook, Instagram and the Meta AI app would follow in the coming days, and the open-weights promise slipped again: August's pledge to release the Muse Spark 1.2 weights became an undated line about open weights coming soon. Full detail in our Muse Spark 1.3 breakdown.
Sep 2 Google Gemini 3.8 Flash, three points on the Intelligence Index at the same $0.75 / $3.75 New
Google made Gemini 3.8 Flash generally available on September 2, 2026, three weeks after 3.7 Flash and its third Flash release in 43 days. Every specification carries over unchanged: 1M input context, a 64K output limit, a March 2026 knowledge cutoff, the same modality list, and the same introductory $0.75 / $3.75 per 1M tokens that doubles to $1.50 / $7.50 on January 1, 2027. Only the scores moved, and they moved on every published row. DeepSWE v1.1 goes from 65.3% to 73.7%, Terminal-bench 2.1 from 85.8% to 89.4%, Terminal-bench 4.0 from 11.2% to 19.1% and BioMysteryBench Human Difficult from 43.5% to 56.5%. Artificial Analysis scores it 59 on Intelligence Index v4.1.1 against 56 for 3.7 Flash, and measures 304.6 output tokens per second against 279.4, with a slow 13.39-second time to first token that makes it a batch and agent model rather than an interactive one. Read the wins alongside the losses: Google's own comparison table gives 3.8 Flash 8 of 14 rows, and Claude Opus 5 still takes Terminal-bench 4.0 by 51.8% to 19.1%, OSWorld-2.0 by 75.4% to 59.0% and GDPVal-AA v2 by 1824 to 1545 on Elo. The DeepSWE figure republished across much of the launch coverage, 71.0%, is not the one in Google's PDF, which reads 73.7% against Opus 5's 74.0%. Developer access was live at announcement through the Gemini API, AI Studio, Antigravity and the Gemini Enterprise Agent Platform. Consumer access is reported for Google AI Pro and Ultra subscribers, but the Gemini Apps release notes carry no entry for it yet, and the free Gemini tier still runs 3.6 Flash. Full detail in our Gemini 3.8 Flash breakdown.
Sep 1 Anthropic Claude Fable 5.1 and Mythos 5.1, a new #1 on Artificial Analysis at 66 and a 75% cut to cache reads New
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, and they are one model in two safeguard configurations rather than two models. Fable 5.1 is generally available; Mythos 5.1 relaxes the biology and cybersecurity restrictions and is handed out by invitation, with no published eligibility criteria. Artificial Analysis scores Fable 5.1 at 66 on its v4.1.1 Intelligence Index at max effort, the highest score it has ever measured, ahead of Claude Opus 5 at 63, Claude Fable 5 at 62, and GPT-5.6 Sol and Grok 4.6 at 61. It also took the then-current Agentic Index at 61 against Opus 5's 59, the GDPval-AA v2 professional-deliverables board at 1853 against 1824, and Humanity's Last Exam at 59.1% against Fable 5's 55.5%. The effort setting matters more than usual here: the same model reads 58 at low, 60 at medium, 62 at high and 65 at xhigh, and those settings span an eleven-fold difference in output tokens, from 13.1 million to 143.7 million across the evaluation suite. Pricing is unchanged at $10 / $50 per 1M tokens, but cache reads drop 75% to $0.25 from Fable 5's $1.00, which is where the saving actually sits on long agentic runs. On Anthropic's own figures it reaches 52.6% on Terminal-Bench-Science 0.1 against Opus 5's 29.0%, and the 212-page system card raised the company's own alignment risk assessment from very low to low. Two things to weigh before switching: Anthropic still recommends starting with Opus 5 for most workloads at half the price, and at launch no Arena board had votes for either model. Fable 5.1 has since been rated: it is second on both of Arena's coding boards and fifth on the text board. Full detail in our Claude Fable 5.1 and Mythos 5.1 breakdown.
Aug 28 Z.ai GLM 5.3 weights are out after the safety hold, but the licence is not MIT New
Z.ai published the GLM 5.3 weights on Hugging Face on August 28, 2026, two weeks after the API launch and on exactly the date its placeholder had carried, so the safety review it promised did happen and the hold is now closed rather than open-ended. What changed is the licence. GLM 5.3 Flash had shipped under plain MIT two days earlier; the flagship carries a bespoke document titled the GLM 5.3 License, and the model card's licence field reads glm-5.3 rather than mit. The grant itself tracks MIT closely, with rights to run, deploy, fine-tune, modify, distribute and sell, and with weights explicitly written into the definition of the software, so commercial use is open to almost everyone. One clause is genuinely new: a model-as-a-service operator whose revenue passes $10 billion over any twelve consecutive months has to pass a Z.ai security review before deploying commercially, and Z.ai reserves the right to decide that review's scope and method itself. Note also that Hugging Face counts the published checkpoint at 753B parameters against the 744B base Z.ai describes it sharing with GLM-5.2, which is the sort of gap a mechanical parameter count produces rather than evidence of a different model. This does not move our open-weight pick. GLM 5.3 Flash stays the model we would host, because 320B under plain MIT is a far easier thing to run and to clear legally than 753B under a bespoke licence. Full detail in our GLM 5.3 breakdown.
Aug 28 Tencent Tencent Hy4 Preview, 770B open weights under Apache 2.0 and the largest permissive model yet New
Tencent released and open-sourced Hy4 preview on August 28, 2026, the new flagship of its Hunyuan line and the largest model anyone has published under Apache 2.0. It runs 770 billion total parameters with 49 billion active per token on a mixture-of-experts design, takes a context window Tencent describes as exceeding 1M tokens, and ships with an FP8 build alongside the full-precision weights. API pricing is $0.834 per 1M input tokens, $2.501 output and $0.042 for cached input, which is inexpensive for the size. The licence is the real headline: Apache 2.0 is more permissive than Kimi K3's custom document and than the bespoke licence Z.ai attached to GLM 5.3 the same day. Treat the quality claim carefully, because the only evidence so far is the vendor's. Tencent ran an internal blind evaluation of 203 engineering tasks scored by 163 experts, where Hy4 preview takes 2.99 out of 4.00 against Kimi K3's 2.94 and GLM 5.3's 2.92. That is Tencent's own panel, the margins sit inside a tenth of a point, and no independent house had published an Intelligence Index or an Arena rating for it at the time. Arena has since placed it eleventh on WebDev at 1624, while Artificial Analysis still has no Intelligence Index for it, so we are listing it rather than ranking it. Tencent also says the model helped automate the optimisation of its own training pipeline and lifted its own inference throughput by a measured 31.8%. It is available as open weights and through WorkBuddy, CodeBuddy, Yuanbao and ima, plus Tencent Cloud TokenHub and OpenRouter, with free access on WorkBuddy and CodeBuddy for two weeks from launch. Full detail in our Tencent Hy4 Preview breakdown.
Aug 26 Z.ai GLM 5.3 Flash, Ox Alpha unmasked, and the first MIT weights Z.ai has shipped since GLM-5.2 New
Z.ai released GLM 5.3 Flash on August 26, 2026, and it is not the model the field was waiting for. The weights held back on August 14 belong to GLM 5.3; this is a separate model on a newly trained base, 320 billion parameters with 18 billion active, the first natively multimodal model in the GLM-5 line, and it went up on Hugging Face under MIT the same day. It is also the stealth model that had been serving traffic on OpenRouter as Ox Alpha, run during that preview entirely on Chinese AI chips, which describes how it was served rather than how it was trained. Artificial Analysis scores it 57 on the Intelligence Index v4.1.1, four points above GLM-5.2 at less than half the size, and places it on the intelligence-versus-cost Pareto frontier. On today's v4.3 index it reads 41.9 against GLM-5.2's 34.0, at $0.25 an index task. List price is $0.15 / $0.50 per 1M tokens; the 50% launch promotion ran to 24:00 on September 9 and has expired. DeepSeek's V4.1-Flash replacement on September 10 closed most of the gap: off-peak the two now tie at $0.15 on input, with GLM cheaper on output at $0.50 against $0.60, while at peak GLM undercuts it on both. Treat the coding and agentic figures with care: they come from Z.ai's own chart, which picks Claude Opus 4.8 and GPT-5.6 Terra as comparators rather than the current leaders, and it reads Terminal Bench 2.1 84.3, DeepSWE v1.1 63.4 against GLM-5.2's 46.2, and AutomationBench 48.8. The one number on that chart Artificial Analysis ran is GDPval-AA v2, where GLM 5.3 Flash takes 1773 Elo, the highest plotted. Arena has since rated it, seventeenth on WebDev and tenth on image-to-WebDev. This does move our open-weight pick: GLM 5.3 Flash replaces GLM-5.2 as the MIT model we would host. One caution on hardware, 18B active does not mean 18B to host, because mixture-of-experts routing needs all 320B weights resident, so this is a multi-GPU server rather than a workstation. Full detail in our GLM 5.3 Flash breakdown.
Aug 26 Alibaba Qwen3.8-Flash-Next, open weights and the first public look at the Qwen4 architecture New
Alibaba published Qwen3.8-Flash-Next on Hugging Face on August 26, 2026, and the point of it is the architecture rather than the scoreboard: Qwen says it is an early look at the design the Qwen4 family will use, released so the community can prepare for it. It is a 125-billion-parameter mixture-of-experts with only 6 billion parameters active per token, plus a separate 51-billion-parameter N-gram embedding, and it takes text, image and video in. Context is 262,144 tokens natively, extensible to 1M. Qwen's published figures include 91.7 on GPQA Diamond, 62.5 on SWE-Bench Pro, 58.7 on DeepSWE 1.1, 81.0 on SWE-Bench Multilingual and 84.5 on AndroidWorld vision, all vendor-reported and, at the time, with no independent rating at all. Artificial Analysis has since scored it 39.9 on Intelligence Index v4.3 and Arena places it ninth on WebDev at 1635. The licence needs reading before you build on it, because it is not Apache. The Qwen Community License 1.0 allows commercial use, modification and hosting, but it requires prominent display of the model name once a product passes 100 million monthly active users or $20 million in monthly revenue, and it requires a separate licence from Qwen to run a model-as-a-service or an AI work-assistant business on it. Internal use is exempt from that second condition.
Aug 21 OpenAI GPT-5.6 Sol cut by more than 20%, now under Claude Opus 5 on both sides New
OpenAI cut GPT-5.6 Sol from $5 / $30 to $4 / $20 per 1M tokens on August 21, 2026, the first cut to the flagship since the GPT-5.6 family launched in July. The rate is promotional and OpenAI says it runs about three months. It covers the API and credits on Codex and ChatGPT Work; Plus, Pro and Business subscription prices are unchanged. At $4 / $20 Sol now undercuts Claude Opus 5 at $5 / $25 on both input and output, the first time OpenAI's flagship has been the cheaper of the two.
Aug 16 DeepSeek DeepSeek raises API prices roughly fourfold and splits the rate card into peak and off-peak tiers New
DeepSeek moved both V4 models onto a new rate card at 16:00 UTC on August 16, 2026, and the direction is unusual for this field: prices went up, by roughly four times at peak. V4-Flash 0731 went from a flat $0.14 / $0.28 per 1M tokens to $0.22 / $0.66 off-peak and $0.44 / $1.32 at peak, and V4-Pro 0813 from $0.435 / $0.87 to $0.66 / $1.98 off-peak and $1.32 / $3.96 at peak, with cache-hit input at $0.007 and $0.022 off-peak respectively. Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, and every other hour is off-peak at exactly half the peak rate. DeepSeek framed the change as allocating capacity more reasonably. Nothing about either model changed, so this is purely an economics event, but it is a consequential one for this page: it cost DeepSeek V4-Flash the price-performance pick, which moved to GLM 5.3 Flash at $0.15 / $0.50 list. DeepSeek reversed most of the rise on September 10, retiring V4-Flash and replacing it with V4.1-Flash at $0.15 / $0.60 off-peak, so the two are level on input off-peak again. The weights are still MIT, so anyone self-hosting is unaffected. Full detail in our DeepSeek V4 breakdown.
Aug 14 Z.ai GLM 5.3, a large agentic jump from post-training alone, shipped without the open weights New
Z.ai released GLM 5.3 on August 14, 2026, and the notable thing is what did not change: it runs on the same 744-billion-parameter base as GLM-5.2, with no new pretrain. Every gain comes from scaled post-training, and the agentic ones are large. On Z.ai's own chart Terminal-Bench 3.0 goes from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and CyberGym from 77.2% to 84.5%, which the company says edges Claude Mythos 5 at 83.8 and GPT-5.6 Sol at 83.6. Every one of those figures is the vendor's, with no independent audit yet, and the chart is less flattering elsewhere: on Z.ai's internal Code Bench, Claude Fable 5 still leads at 39.5% against GLM 5.3's 31.4%, and on ExploitBench the frontier leaders sit at 78.0% against 54.4%. The catch is the licence, not the score. GLM-5.2 shipped MIT weights within days; GLM 5.3 shipped with none, and Z.ai said they would follow after a safety evaluation of the model's vulnerability-discovery capability, roughly two weeks out. That resolved on August 28, 2026, when Z.ai published the weights on exactly the date the placeholder had carried, though under a bespoke GLM 5.3 License rather than the plain MIT that GLM-5.2 and GLM 5.3 Flash both got. Per-token access is now published at $1.40 / $4.40 per 1M tokens, alongside the GLM Coding Plan and ZCode. What arrived on August 26 instead was GLM 5.3 Flash, a different and smaller model that did ship MIT weights, and that is the release that moved our open-weight pick. Full detail in our GLM 5.3 breakdown.
Aug 13 Google Gemini 3.7 Flash, coding scores jump and the price halves to $0.75 / $3.75, until January New
Google shipped Gemini 3.7 Flash on August 13, 2026, 23 days after 3.6 Flash and its third Flash release in under two months, while Gemini 3.5 Pro remains unreleased. The coding gains are the point: DeepSWE v1.1 climbs from 49.0% to 65.3%, FrontierCode 1.1 Main from 34.4% to 43.6%, and AutomationBench nearly doubles from 17.0% to 30.4%. Artificial Analysis scores it 56 on its v4.1.1 Intelligence Index against 52 for 3.6 Flash, and ranks it first of 182 models on output speed at 385.1 tokens per second, which puts it on the intelligence-versus-time Pareto frontier. The context window and 64K output limit are unchanged from 3.6 Flash, so nothing was traded away. Two caveats carry more weight than the benchmarks. The $0.75 / $3.75 rate is introductory and expires on December 31, 2026, after which it doubles to $1.50 / $7.50, exactly what 3.6 Flash cost at launch; Google also moved 3.6 Flash onto the same rate, so the two now cost the same and the upgrade is a pure capability decision. And consumer access runs only through Spark, which needs Google AI Pro or Ultra and excludes the European Economic Area, the UK, Switzerland and Nigeria, so most European readers cannot reach it in the Gemini app at all. The free Gemini tier still gets 3.6 Flash. API access is unaffected everywhere. Full detail in our Gemini 3.7 Flash breakdown.
Aug 12 SpaceXAI Grok 4.6, post-training refresh takes Intelligence Index 56 to 61 at an unchanged $2 / $6 New
SpaceXAI released Grok 4.6 on August 12, 2026, and it is explicitly not a new foundation model: the company kept the Grok 4.5 base and bought the gains with a longer supplemental training run, regenerated fine-tuning trajectories and reinforcement learning inside agentic environments. No new architecture or parameter count was published. Artificial Analysis scores it at Intelligence Index 61, five points above Grok 4.5, which ties GPT-5.6 Sol for joint third behind Claude Opus 5 at 63 and Claude Fable 5 at 62 on the index as it stood that month; Claude Fable 5.1 has since gone in above all three at 66. The agentic figures are stronger than the composite: a GDPval-AA v2 Elo of 1753 behind only Opus 5, 88.4% on Terminal-Bench v2.1 and 50.7% on tau3-Banking. Efficiency was the argument at the time, at $0.84 per index task and roughly 53 turns and 0.5 billion input tokens on long-horizon work against about 103 turns and 2.0 billion for Opus 5. On Artificial Analysis's current v4.3 figures that argument has weakened: Grok 4.6 costs $1.86 an index task, more than Muse Spark 1.3 at $1.61 for a higher score. Pricing is unchanged at $2 / $6 per 1M tokens with cached input at $0.50, on the same 500K-token context window, but there is a trap worth knowing: cross 200,000 prompt tokens and SpaceXAI re-bills the entire request at $4 / $12, not just the overage. One warning on the benchmark coverage, because a lot of reporting has mangled it: Terminal-Bench v3.0 is a different and much harder test than v2.1, so Grok 4.6 scoring 26% on v3.0 and 88.4% on v2.1 are both true of the same model. It went live day one in the SpaceXAI API, Cursor and Grok Build, plus OpenRouter, Vercel and Cloudflare. Full detail in our Grok 4.6 benchmarks and pricing breakdown.
Aug 5 Meta Muse Spark 1.2 and Muse Code, Intelligence Index 51 to 57 at an unchanged $1.25 / $4.25, plus a terminal coding agent New
Meta shipped Muse Spark 1.2 on August 5, 2026 alongside Muse Code, its first terminal coding agent, which runs on macOS and Linux with no desktop app and no subscription. Artificial Analysis scores the model at Intelligence Index 57 at its highest reasoning setting on its current v4.1.1 index, up from 51 for 1.1 and 43 for the April release, and almost all of that gain is agentic rather than general knowledge; its GDPval-AA v2 Elo jumped from 1371 to 1631 while Terminal-Bench v2.1 moved only from 78% to 80%. Pricing is unchanged at $1.25 / $4.25 per 1M tokens on a 1M-token context window, and Meta added a Contributor tier at $0.10 / $0.20, roughly 92% cheaper, in exchange for letting Meta train on your prompts and completions. Availability widened from the US-only preview that 1.1 shipped under to expanded global access across Muse Code, the Meta Model API and OpenRouter. On Meta's own launch charts the model finishes second to Claude Opus 5 on all three coding benchmarks it published, at 82.9% against 86.7% on Terminal-Bench 2.1.
Aug 3 Alibaba Qwen 3.8 (Qwen3.8-Max), 2.4-trillion-parameter multimodal flagship now generally available at $2 / $6 New
Alibaba made Qwen3.8-Max generally available on August 3, 2026, two weeks after previewing it at WAIC Shanghai, and every figure that was missing at preview now has a number behind it. It runs 2.4 trillion total parameters with roughly 95 billion active per token on a sparse Mixture-of-Experts design, takes text, images and video, and confirms a 1M-token context window with up to 128K output tokens. API pricing is $2 / $6 per 1M tokens, flat across the entire context rather than tiered by prompt length, with cached reads at $0.25. The "second only to Fable 5" claim now splits cleanly in two: on Arena's vision board it is #2 at 1305, behind only Fable 5, but on the text board it is #5 at 1496, behind four Anthropic entries. Both Arena scores are still flagged preliminary, and the promised open weights have not shipped. We are keeping Qwen 3.8 out of the ranked picks until the weights and the licence actually land.
Pending Google Gemini 3.5 Pro, still unreleased, months behind schedule per Bloomberg Delayed
Gemini 3.5 Pro remains the biggest pending launch. Google announced it at I/O on May 19 alongside Gemini 3.5 Flash, but only Flash shipped, and the June target slipped. Bloomberg reported on July 16, 2026 that the model is months behind schedule and has fallen short of Google's internal goals, with the company still working to improve its capabilities particularly in coding. That is the whole sourced picture. Google has never publicly confirmed a launch date, and Google has not published final specs such as the context window or reasoning modes, so treat circulating figures as unconfirmed. Use Gemini 3.8 Flash in the meantime if you are working through the API, or Gemini 3.6 Flash if you are on the free Gemini app, which still runs it; we will move 3.5 Pro into the main ranking the moment it goes live.
Category pick, September 2026

Best AI for Writing

1
Claude Fable 5.1
Best writing overall
2
Kimi K3
Creative fiction
3
Claude Sonnet 5
Free everyday

The best AI for writing is Claude Fable 5.1, which beats Claude Fable 5 on every board that has measured both, at the same $10 / $50, with Claude Sonnet 5 as the free-tier value pick and Claude Fable 5 as the pick if you weight Arena's human votes above the harness scores. Fable 5 leads Arena's creative-writing leaderboard and tops LiveBench Language at 90.7, but it has slipped to eighth on EQ-Bench Creative Writing v3, a board now led by GPT-6 Astra at 2163.9 with Claude Fable 5.1 second at 2152.7. This is a change from last month, when Claude Sonnet 5 held this slot on GDPval-AA; the preference boards do not support that placing, so Sonnet 5 is now the value pick rather than the quality leader. Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. If your writing is a work deliverable rather than prose, the GDPval-AA v2 professional-deliverables board is now led by Claude Fable 5.1 at 1853, ahead of Claude Opus 5 at 1824 and Fable 5 at 1723. Fable 5.1 landed on September 1 as the higher-capability model at the same $10 / $50, and it beats Fable 5 on every board that has measured both, including Humanity's Last Exam at 59.1% against 55.5%. We have moved the crown to it, because this page ranks on measured scores; Arena has not rated it yet, which is why Fable 5 stays in the table as the human-preference pick. Translation is a separate question with a separate winner, which we work through in our guide to the best AI for translation.

ModelBest ForStrengthWeaknessPrice (per 1M tokens)
Claude Fable 5.1Best writing overall#1 Humanity's Last Exam (59%), #1 Intelligence Index (53.4, v4.3), #1 AA-Omniscience accuracy (67%)Fifth on Arena's text board, behind Claude Fable 5$10 / $50
Claude Fable 5Top of Arena's human-vote boards#1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #8 EQ-BenchPriciest option here$10 / $50
Kimi K3Creative fiction and voice#4 EQ-Bench Creative Writing (2070.6)Only #10 on Arena creative writing$3 / $15
Claude Sonnet 5Free everyday writingFree and default on claude.ai, 1M context#53 Arena creative writing; 75.0 LiveBench Language$2 / $10, now permanent
Claude Opus 5Professional deliverables on a budget#2 GDPval-AA v2 at 1824, ahead of Fable 5 (1723)Behind Fable 5 on Arena text, behind Fable 5.1 on GDPval-AA v2$5 / $25
GPT-5.5Fact-anchored business writingDocumented factual-reliability gains over GPT-5.4Reasoning tiers now marked deprecated$5 / $30
Gemini 3.8 FlashBulk drafts at scaleIntelligence Index 41.2, 304.6 tok/s output, 1M contextAgent-tuned rather than prose-tuned; intro price doubles January 1, 2027$0.75 / $3.75
Runner-up and alternatives: Claude Fable 5 is the runner-up and the pick if you weight Arena's human votes, Kimi K3 is the pick for distinctive fiction though it now sits fourth on EQ-Bench, Claude Sonnet 5 is the value pick and the one to use if you are not paying, Claude Opus 5 is the cheaper pick for professional deliverables, and Gemini 3.8 Flash is the pick for bulk drafting.
What changed this month

The writing crown moved to Claude Fable 5.1, and the rule behind this page moved with it: we now rank on measured scores rather than waiting for Arena's human votes. Fable 5.1 beats Fable 5 on every board that has measured both, taking Humanity's Last Exam at 59.1% against 55.5%, GDPval-AA v2 at 1724 against 1632 and the Intelligence Index at 53.4 against 49.7 on v4.3, all at the same $10 / $50. Claude Fable 5 keeps Arena's text and creative-writing boards, where it is still #1 at 1508.6 on the August 1 cutoff, so it stays in the table for anyone who weights human preference above harness scores, and Claude Sonnet 5 remains the free value pick. Two numbers we quoted last month were also re-fitted: AA-Omniscience now reads 43 for both Fable models rather than 40, and Fable 5 on Humanity's Last Exam reads 55.5% rather than 53.3%.

Category pick, September 2026

Best AI for Chat & Daily Assistant

1
GPT-5.6
ChatGPT's default
2
Claude Fable 5
Highest-rated
3
Claude Opus 5
Thoughtful depth

The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena's text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $4 / $20 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.

ModelBest ForStrengthWeaknessPrice
GPT-5.6Everyday chat, ChatGPT's defaultThe assistant most people can open; Terra matches GPT-5.5 at ~half cost#14 on Arena text; scheming flagged by METRFree / $20/mo Plus; API $0.20 / $1.20 to $4 / $20
Claude Fable 5Highest-rated conversation#1 on Arena text overall (1508.6) and 6 of 7 subcategoriesNo free tier; usage-credit access on Pro$10 / $50 API
Claude Opus 5Thoughtful, nuanced answersThird on Artificial Analysis's Intelligence Index (50.7), behind Claude Fable 5.1 and GPT-6 Astra#10 on Arena text, behind Claude Fable 5$20/mo Pro, $5 / $25 API
GPT-5.5 InstantHallucination-sensitive daily work52.5% fewer hallucinated claims vs 5.3 InstantReasoning tiers now marked deprecated$20/mo Plus; API $5 / $30
Gemini 3.6 FlashFast, free, multimodalStill the model the free Gemini tier serves, 1M context, #12 on Arena textWeaker on hardest reasoning; 3.8 Flash outscores it at the same priceFree / $0.75 / $3.75 API
Fello AIAll the top models, one appChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPadRouted via app, not direct$9.99/mo
Runner-up and alternatives: Claude Fable 5 is the runner-up and the actual preference leader, Claude Opus 5 is the runner-up for thoughtful daily use, Gemini 3.6 Flash is the runner-up for fast and free, and Grok 4.6 is the niche pick for live-news days. Fello AI is the natural pick if you want the top models in one Mac and iOS app for $9.99/month instead of juggling subscriptions.
What changed this month

GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open. OpenAI's GPT-6 Astra launched on September 3 and reached ChatGPT Business and Pro on September 4 and Plus a few hours later, but the pick stays with GPT-5.6 on reach: Astra went first to approved defenders in the Daybreak program, and OpenAI has announced nothing for the free tier.

Category pick, September 2026

Best AI for Images

1
ChatGPT Images 2.5
Text in images
2
MAI-Image-2.6
#1 AA image editing
3
Muse Image
Meta AI ecosystem

The best AI for image generation is ChatGPT Images 2.5, which OpenAI made the default on every ChatGPT tier on September 8 and which took both of Arena's image boards within a day. Its two models finished first and second on each: GPT-Image-2.5 Sunburst, built for precision, leads text-to-image at 1420.7 Elo and image editing at 1520.4, the faster GPT-Image-2.5 Flare is second at 1399.0 and 1490.6, and last month's GPT Image 2 is pushed into third on both. Those placements rest on a few thousand votes each against GPT Image 2's 79,000 and 236,000, and on text-to-image the two share a rank band, so read the gap between Sunburst and Flare as provisional. Artificial Analysis has rated neither model, and its two boards no longer agree with Arena's: GPT Image 2 still leads text-to-image there at 1178.1, but Microsoft's MAI-Image-2.6 has taken image editing at 1122.5. That makes MAI-Image-2.6 the runner-up overall, the highest-placed non-OpenAI model on Arena's text-to-image board and first or second on both Artificial Analysis boards, with Meta's Muse Image third. Reve 2.1 leaves the podium: it is still third on Artificial Analysis's text-to-image board but has fallen to sixth on Arena's and to twelfth and fourteenth on the two editing boards. Our full breakdown of what shipped is in ChatGPT Images 2.5.

ModelBest ForStrengthWeaknessPrice
ChatGPT Images 2.5Images with readable text#1 and #2 on Arena text-to-image (1420.7 / 1399.0) and image editing (1520.4 / 1490.6)Not yet rated by Artificial AnalysisDefault on every ChatGPT tier
MAI-Image-2.6Editing an image you already have#1 on Artificial Analysis image editing (1122.5) and #2 on its text-to-image board (1149.4)Public preview, not in Copilot or Bing yetMAI Playground / Microsoft Foundry
Muse ImageImage editing, Meta ecosystem#3 on Artificial Analysis image editing (1105.1), #5 on its text-to-image boardNew, thin tooling around itMeta AI app
Nano Banana 2 (Gemini 3.1 Flash Image)Photoreal portraits and products#4 on Artificial Analysis text-to-image, ahead of Nano Banana Pro on Arena'sNano Banana Pro is ahead on Arena image editingGemini app / AI Studio
Seedream 5.0 ProMultilingual text + region editing10+ languages incl. Arabic RTL, lasso and layer editingAround tenth on the independent boards; copyright cloudBytePlus / Magnific
Midjourney v8Stylized art, illustrationAesthetic baseline most artists preferWeaker on text in image$10-$120/mo
Grok ImagineNSFW / Spicy ModeMost permissive guardrails; Arena rates its Image 2.0 model #5 on text-to-image and #4 on image editingImage 2.0 is not on Artificial Analysis's boards$30/mo SuperGrok
Runner-up and alternatives: MAI-Image-2.6 is the runner-up overall and the pick for editing an image you already have, Muse Image is third and the pick inside Meta's apps, and Nano Banana 2 is still the photoreal pick. Reve 2.1 remains the pick for layout, typography and native 4K even after leaving the podium. Grok Imagine is still the only frontier model that allows Spicy Mode adult content, and Arena now rates its Image 2.0 model fifth on text-to-image and fourth on image editing.
What changed this month

The image crown moved. OpenAI shipped ChatGPT Images 2.5 on September 8 and its two API models went straight past GPT Image 2 on both of Arena's boards: Sunburst first on text-to-image at 1420.7 and on image editing at 1520.4, Flare second at 1399.0 and 1490.6, GPT Image 2 third on both. The three boards we track no longer agree. Artificial Analysis has not rated the 2.5 models at all, it still has GPT Image 2 first on text-to-image at 1178.1, and on its image-editing board Microsoft's MAI-Image-2.6 has taken first place at 1122.5. MAI-Image-2.6 replaces Reve 2.1 as the runner-up, and Reve 2.1 leaves the podium entirely after falling to sixth on Arena text-to-image.

Category pick, September 2026

Best AI for Video

1
Gemini Omni 1.1 Flash
#1 Arena text-to-video
2
MiniMax H3
2K + native audio
3
Dreamina Seedance 2.0
Closest challenger

The best AI for video generation is Gemini Omni 1.1 Flash, which Google shipped on August 27 and which now leads Arena's text-to-video board at 1515, just ahead of its own predecessor Gemini Omni Flash at 1511. It is the more capable model on specification too: scene extension in 10-second increments to a cumulative 40 seconds, first and last frame control, video references, 360p drafts and 1080p or 4K output, priced at roughly $0.03 per second at 360p, $0.10 at 720p, $0.15 at 1080p and $0.30 at 4K. Gemini Omni Flash remains the cheaper option at about $0.10 a second and is a statistical tie on Arena, but it caps at 10-second clips. Neither leads Artificial Analysis's video board any more: Alibaba's Wan 3.0, an API-only model released on August 24 that builds video from documents, spreadsheets and slides, tops it at 1239 with Omni Flash on 1238 and MiniMax H3 Max on 1235. If you need longer takes inside Google's tooling, Veo 3.1 still runs in the Gemini app, AI Studio and Vertex AI with native audio and 1080p output. MiniMax H3 (July 31) remains the specification contender, with 2K clips of 4 to 15 seconds and native stereo audio from $0.13 a second.

ModelBest ForStrengthWeaknessPrice
Gemini Omni 1.1 FlashBest AI video overall#1 Arena text-to-video (1515); 40-second scene extension, 4K outputSecond to Wan 3.0 on Artificial Analysis's board$0.03-$0.30/sec; Gemini API / AI Studio / Flow
Gemini Omni FlashCheaper 10-second clips#2 Arena text-to-video (1511), #2 AA video (1238); conversational editingCaps at 10-second generations~$0.10/sec; Gemini app / AI Studio
Wan 3.0Video from documents and slides#1 on Artificial Analysis's video board (1239); 2-30s, up to 1080p, Omni-Reference inputAPI-only, no published weights; #3 on Arena$0.05-$0.20/sec
MiniMax H32K clips with native audio4-15s at 2K, native stereo audio; #8 Arena text-to-video (1462)#4 on AA video (1227); the open weights exclude the 2K upscaler and need an application in the US, EU, UK and South Korea$0.13/sec at 2K
Dreamina Seedance 2.5ByteDance challenger#6 Arena text-to-video (1482); leads image-to-video with audioByteDance ecosystem, limited Western accessDreamina / BytePlus
Muse VideoMeta ecosystem video#3 text-to-video at 1459Newest of the group, thin toolingMeta AI app
Veo 3.1Longer production clipsNative audio, 1080p, strong physics consistency#6 on Arena video, not the quality leaderGoogle AI Pro / Ultra
Kling 3.0 / 3.0 TurboFast iteration at lower costNative 4K, 60fps, 15-second clips; Turbo shipped June 17Outside the top 16 on Arena text-to-videoFrom $10/mo
Luma Ray 3Photoreal scenesStrong realism for landscapesSmaller communityFree / from $9.99/mo
Runner-up and alternatives: Dreamina Seedance 2.0 is the runner-up overall and beats Omni Flash on image-to-video with audio, where it leads at 1198 to Omni Flash's 1191, though Omni Flash still leads the no-audio image-to-video split. Muse Video is third, and Veo 3.1 is the pick when 10 seconds is not enough. OpenAI retired the Sora 2 consumer app on April 26, 2026 and only the developer API remains, through September 24, 2026. We track the retirement dates for older models across every provider in a separate running list.
What changed this month

The video crown moved to Gemini Omni 1.1 Flash, which Google shipped on August 27 and which now leads Arena's text-to-video board at 1515 against Gemini Omni Flash's 1511. The two sit inside each other's confidence intervals, so read that as a tie broken on capability rather than a decisive win: 1.1 Flash extends scenes to a cumulative 40 seconds and outputs 4K, where Omni Flash caps at 10 seconds. The bigger change is that neither model leads Artificial Analysis's video board any more. Alibaba's Wan 3.0 took it on August 24 at 1239, ahead of Omni Flash at 1238 and MiniMax H3 Max at 1235, and it accepts documents, spreadsheets and slides as input. Wan 3.0 is API-only with no published weights, so it does not change our open-weight picks. MiniMax H3 also picked up Arena votes this month, entering the text-to-video board at #8 with 1462.

Category pick, September 2026

Best AI for Coding

1
GPT-6 Astra
Leads every coding board
2
Claude Fable 5.1
Highest benchmark scores
3
Claude Opus 5
Best value, $5 / $25

The best AI for coding is GPT-6 Astra, which leads Artificial Analysis's own Terminal-Bench 4.0 at 60% against Claude Fable 5.1's 55% and tops both of Arena's coding boards, WebDev at 1800 and image-to-WebDev at 1733. Claude Fable 5.1 is the runner-up and keeps the composite measures: it leads the Intelligence Index at 53.4 on v4.3 and AA-Briefcase at 1646 against Astra's 1562, the agentic board that replaced the retired Agentic Index. Claude Opus 5 is the value pick, at $5 / $25 per 1M tokens against $10 / $50 for both models above it, taking 49% on Terminal-Bench 4.0 and third place on both Arena coding boards, and Anthropic's own docs still tell developers to start with Opus 5 for complex agentic coding. Read that price argument carefully, because it is about tokens rather than tasks: Opus 5 is half Astra's list price but costs more per Intelligence Index task, $5.86 against $3.26, because it spends more tokens getting there. Alibaba's Qwen3.8-Max-0902 held Arena's WebDev board for about three days in early September and now sits fourth at 1681. Among open weights the evidence splits: GLM-5.3 is the strongest on the harness at 42% on Terminal-Bench 4.0, while Kimi K3 is the highest-placed open model on Arena WebDev, fifth at 1674, but manages only 13% on that harness, and self-hosting it means 1.56 TB of weights. Artificial Analysis has retired both its Coding Index and its Coding Agent Index, so the two composite coding rankings this page used to quote no longer exist. The cheapest serious contender is GLM 5.3 Flash at $0.15 / $0.50 under MIT, which takes 33% on Terminal-Bench 4.0 at two cents a task, with DeepSeek V4.1-Flash level with it on off-peak input since September 10.

ModelBest ForStrengthWeaknessPrice (per 1M tokens)
GPT-6 AstraBest coding overall60% Terminal-Bench 4.0 (xhigh), #1 Arena WebDev (1800) and #1 image-to-WebDev (1733)Twice Opus 5's price; behind Claude Fable 5.1 on the Intelligence Index and AA-Briefcase$10 / $50
Claude Fable 5.1Highest composite scores#1 Intelligence Index (53.4, v4.3) and #1 AA-Briefcase (1646); 55% Terminal-Bench 4.0Second to Astra on all three coding boards; twice the price of Opus 5$10 / $50
Claude Opus 5Best value at half the price49% Terminal-Bench 4.0 and #3 on both Arena coding boards; Anthropic's docs still say to start hereCosts more per index task than Astra, $5.86 against $3.26$5 / $25
Qwen3.8-Max-0902Brief holder of Arena's WebDev board#4 Code Arena: WebDev (1681); Intelligence Index 45.4; 2.4T parameters, 1M contextQwenCloud API only with no consumer front-end; lost the WebDev lead within days$2 / $6
Claude Fable 5Hardest long-horizon agentic workIntelligence Index 49.7; 42% Terminal-Bench 4.0; 1M contextThe priciest per index task on this table at $8.75; #6 on Arena image-to-WebDev$10 / $50
Kimi K3Highest-placed open model on Arena#5 Arena WebDev (1674), the best open placement on that board; 2.8T/104B activeOnly 13% on Terminal-Bench 4.0; 1.56 TB to self-host$3 / $15
GPT-5.6 SolOpenAI's cheaper flagshipIntelligence Index 47.1; 40% Terminal-Bench 4.0Well behind Astra on every coding board; eval-gaming flagged by METR$4 / $20
Muse Spark 1.3Cheap agentic codingIntelligence Index 48.2 at $1.25 / $4.25, and $1.61 per index task33% on Terminal-Bench 4.0; the coding wins come from Meta's own chart, measured on a partners-only max variant$1.25 / $4.25
Grok 4.6Cheap value coder88.4% on Terminal-Bench v2.1; Intelligence Index 44.421% on Terminal-Bench 4.0; prompts of 200K tokens and up re-bill the whole request at double$2 / $6
Gemini 3.8 FlashAgent coding at scaleIntelligence Index 41.2; 304.6 output tokens per second20% on Terminal-Bench 4.0; intro price doubles January 1, 2027$0.75 / $3.75
GLM 5.3 FlashCheapest serious coder you can host33% Terminal-Bench 4.0 at two cents a task; MIT, 320B/18B activeGLM-5.3 reaches 42% on the same harness; no consumer productOpen weights (MIT)
Runner-up and alternatives: Claude Fable 5.1 is the runner-up and still tops the Intelligence Index and AA-Briefcase, Claude Opus 5 is the value pick at half the list price of either model above it, Kimi K3 is the highest-placed open model on Arena's WebDev board while GLM-5.3 is the strongest open model on Artificial Analysis's Terminal-Bench 4.0, and GLM 5.3 Flash is the pick for teams hosting their own, at two cents a task under plain MIT. Inside IDEs, Cursor with Claude is still the most popular pairing and Claude Code is the natural pick if you live in the terminal.
What changed this month

The coding crown moved to GPT-6 Astra. Artificial Analysis retired both its Coding Index and its Coding Agent Index, which is where Claude Fable 5.1's lead lived, and on the boards that remain Astra is ahead on all three: 60% against 55% on Artificial Analysis's own Terminal-Bench 4.0, 1800 against 1758 on Arena's WebDev board, and 1733 against 1710 on image-to-WebDev. Claude Fable 5.1 keeps the composite measures, the Intelligence Index at 53.4 and AA-Briefcase at 1646 against Astra's 1562, and becomes the runner-up. Claude Opus 5 stays the value pick at $5 / $25, though that case is now about token price rather than task cost: it costs $5.86 an Intelligence Index task against Astra's $3.26. Alibaba's Qwen3.8-Max-0902 held WebDev for about three days and is now fourth. Among open weights the evidence has split, with Kimi K3 highest on Arena and GLM-5.3 far ahead on the harness, 42% against 13%.

Category pick, September 2026

Best AI for Creativity

1
Grok 4.6
Fewest guardrails
2
Claude Fable 5
Highest-quality prose
3
Kimi K3
Fiction and voice

The best AI for unfiltered, on-trend creative work is Grok 4.6, and we want to be exact about why. This pick is about the product, not the prose quality. Grok carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, and the boards were blunt about its predecessor: Grok 4.5 sat #33 on EQ-Bench Creative Writing and #41 on Arena's creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. Neither creative-writing board has rated Grok 4.6 yet, and since SpaceXAI aimed this release at coding and agentic work rather than prose, we are not assuming it moved. If you are picking on output quality alone, Claude Fable 5 still leads Arena's creative-writing board, while EQ-Bench now puts GPT-6 Astra first at 2163.9, Claude Fable 5.1 second and Kimi K3 fourth. Choose Grok 4.6 for what it will let you make, not for how well it writes.

ModelBest ForStrengthWeaknessPrice
Grok 4.6Unfiltered, opinionated, on-trendFewest content restrictions, native real-time X groundingCreative-writing boards have not rated it; Grok 4.5 sat #33 EQ-Bench, #41 Arena$30/mo SuperGrok
Claude Fable 5Highest-quality creative prose#1 Arena creative writing, #1 LiveBench LanguageCautious guardrails on edgy briefs$10 / $50 API
Kimi K3Fiction and distinctive voice#4 EQ-Bench Creative Writing at 2070.6Only #10 on Arena creative writing$3 / $15
Claude Opus 5Long-form structured creativityHolds long threads and self-edits; Intelligence Index 50.7, behind Claude Fable 5.1 and GPT-6 AstraMost cautious of the group$20/mo Pro; $5 / $25 API
Gemini 3.1 ProMultimodal creativeStrong text, image and video chainQuotas inside the Gemini appFree / $2.00-$4.00 API in
Grok Imagine (Spicy Mode)NSFW / adult creativeMost permissive image generationNiche use case$30/mo SuperGrok
Runner-up and alternatives: Claude Fable 5 is the runner-up and the right pick if quality matters more than freedom, Kimi K3 is the pick for fiction, and Claude Opus 5 is the pick for creative projects that run across many turns. For adult creative work, Grok Imagine Spicy Mode is still the only frontier-grade option.
What changed this month

Grok keeps this pick, now on Grok 4.6, which replaced Grok 4.5 as SpaceXAI's flagship on August 12. Nothing about the product's permissiveness or its live X access changed, which is what the pick rests on. The model underneath is meaningfully stronger, up five points on the Intelligence Index, 44.4 against Grok 4.5's 39.1 on v4.3, but that gain landed in coding and agentic work rather than prose, and no creative-writing board has rated 4.6 yet. Claude Fable 5 remains the model to use when you want the better writing.

Category pick, September 2026

Best AI for Accuracy & Research

1
Claude Fable 5.1
Highest factual accuracy
2
Gemini 3.1 Pro
Best value, $0.52/task
3
GPT-5.6 Sol
Novel reasoning

The best AI for accuracy and research is Claude Fable 5.1, which holds the highest factual accuracy Artificial Analysis has measured, 67% on AA-Omniscience, and tops Humanity's Last Exam at 59.1%. The value pick is Gemini 3.1 Pro, and it is still what we would reach for when cost matters. Its strongest result is on ARC Prize's ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you want when the answer has to be current rather than merely plausible. It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity's Last Exam, and scores 46.44 on Scale SEAL's HLE board, which Claude Fable 5 now leads at 55.5%. We have dropped the previous ARC-AGI-2 framing: its 77.1% is still correct, but the board has moved and that score now places it around 14th. Two honest caveats. On grounded search specifically, Arena's search leaderboard is led by OpenAI rather than by Google or Anthropic: GPT-5.6 Sol tops it at 1257, with Claude Fable 5 fifth and Gemini 3.1 Pro grounding ninth. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience, a board Claude Fable 5.1 now shares the top of.

ModelBest ForKey BenchmarkWeaknessPrice
Claude Fable 5.1Highest measured factual accuracy67% factual accuracy on AA-Omniscience, the highest Artificial Analysis has measured; #1 Humanity's Last Exam (59.1%)No cheap tier, and unrated on Arena's search leaderboard$10 / $50
Gemini 3.1 ProBest value, cheap factual work98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol)ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search$2.00-$4.00 / $12.00-$18.00 (tiered)
Claude Fable 5Grounded search#5 on Arena's search leaderboard, behind GPT-5.6 SolNo single cheap tier$10 / $50 (Fable 5)
GPT-5.6 SolNovel reasoning#1 ARC-AGI-2 at 92.5%, against a 100% human panelScheming flagged by METR$4 / $20
Claude Opus 5Hardest unseen problems#1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize)Artificial Analysis measures a 50% hallucination rate$5 / $25
Qwen 3.7 MaxFrontier accuracy at value pricing92.4 GPQA Diamond, 200 free requests/dayAPI-only, no chat front-end$1.25 / $3.75 promo; $2.50 / $7.50 list
Claude Opus 4.6Honesty under pressure#1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5Superseded as a flagshipLegacy Anthropic model
Runner-up and alternatives: Gemini 3.1 Pro is the runner-up and the value pick at $0.52 a task with Google Search grounding, GPT-5.6 Sol is the runner-up for novel reasoning, Claude Opus 4.6 sweeps the honesty-under-pressure board, and Qwen 3.7 Max is the value pick at the frontier.
What changed this month

The accuracy crown moved to Claude Fable 5.1, which holds the highest factual accuracy Artificial Analysis has measured, 67% on AA-Omniscience against Claude Fable 5's 65%, and tops Humanity's Last Exam at 59.1% against 55.5%. The AA-Omniscience index itself reads 43 for both models, up from the 40 we quoted last month, because 5.1 buys its extra accuracy with a higher attempt rate on questions it cannot answer. Gemini 3.1 Pro becomes the value pick and is still what we would reach for on cost: 98% on ARC-AGI-1 at $0.52 a task, tying the human panel, plus native Google Search grounding, though its GPQA Diamond re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly. GPT-6 Astra (September 3) is the launch to watch here rather than a new crown: Artificial Analysis measures its hallucination rate falling from Sol's 92% to 51% at max effort while accuracy rises four points, the largest honesty improvement any vendor has posted this year, but it loses about 80 Elo on GDPval-AA v2 and is not on ChatGPT's free tier.

Category pick, September 2026

Best AI for Problem Solving

1
GPT-6 Astra
Reasoning leader
2
Claude Opus 5
Agentic chains
3
GPT-5.6 Sol
Value STEM flagship

The best AI for hard problem solving is GPT-6 Astra, which took the reasoning boards off GPT-5.6 Sol within days of launching. It leads LiveBench Reasoning at 92.65 and ARC-AGI-2 at 95%, the highest anyone has scored against that board's 100% human panel, and sits second on LiveBench Mathematics at 96.81 behind Claude Fable 5.1's 97.01. It also reports 97.6% on FrontierMath Tier 4 v2, a number worth carrying with Epoch AI's disclosure that OpenAI funded the benchmark and holds exclusive access to part of it. The catch is price and reach: $10 / $50 per 1M tokens against Sol's $4 / $20, and it needs a paid ChatGPT plan. GPT-5.6 Sol is the value alternative, third on both LiveBench boards at 96.20 and 91.65 and second on ARC-AGI-2. Qwen 3.7 Max is the budget pick for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, with 200 free model requests a day. Claude Opus 5 remains the alternative for long agentic reasoning chains, though Astra has taken ARC-AGI-3 from it as well.

ModelBest ForKey BenchmarkWeaknessPrice
GPT-6 AstraHardest math, science and reasoning#1 LiveBench Reasoning (92.65), #1 ARC-AGI-2 (95%), #1 ARC-AGI-3Needs a paid ChatGPT plan; 2.5x Sol's price$10 / $50
Claude Opus 5Long agentic reasoning chains#2 AA-Briefcase (1645) and #2 GDPval-AA v2 (1737); 30.2% ARC-AGI-3Astra now leads ARC-AGI-3; 50% hallucination rate$5 / $25
GPT-5.6 SolValue STEM flagship#2 ARC-AGI-2 (92.5%); #3 LiveBench Mathematics (96.20) and Reasoning (91.65)FrontierMath still unpublished; scheming flagged by METR$100/mo ChatGPT Pro; API $4 / $20
Qwen 3.7 MaxCompetition math on a budget97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/dayAPI-only$1.25 / $3.75 promo; $2.50 / $7.50 list
Claude Fable 5Math inside a coding workflow#1 on Arena's math subcategory (1543), 96.0 LiveBench MathematicsPriciest option here$10 / $50
GLM 5.3 FlashOpen-weight problem solvingIntelligence Index 41.9 under MIT, almost eight points above GLM-5.2 at less than half the size; 320B/18B active, 1M contextNeeds a multi-GPU server, not a workstation; GLM 5.3 scores higher on Z.ai's own figures and is downloadable since August 28, but at more than twice the size and under a bespoke licenceOpen weights (MIT)
Runner-up and alternatives: GPT-5.6 Sol is the runner-up and the value pick at $4 / $20, Claude Opus 5 is the natural pick for long-chain agentic reasoning, Qwen 3.7 Max is the budget pick, and GLM 5.3 Flash is the open-weight pick. Our dedicated guide to the best AI for math covers the task-by-task split and the free options.
What changed this month

The problem-solving crown moved to GPT-6 Astra, which took the boards this category is decided on within days of its September 3 launch. It leads LiveBench Reasoning at 92.65 against GPT-5.6 Sol's 91.65, takes ARC-AGI-2 at 95% against Sol's 92.5%, and displaced Claude Opus 5 on ARC-AGI-3, where its 62.7% on the standard harness beats Opus 5's 30.2%. Read the ARC-AGI-3 figures carefully: that board measures human-level action efficiency on unseen environments rather than problems solved, and the widely quoted 98.6% comes from a provider-adapter harness rather than the standard one. Sol drops to third on both LiveBench boards but stays the value alternative at $4 / $20 against Astra's $10 / $50, and Claude Fable 5.1 quietly took LiveBench Mathematics at 97.01.

Category pick, September 2026

Best AI Agent

1
Gemini Spark
Cloud-resident
2
Claude Cowork
Desktop-resident
3
ChatGPT Codex
Coding agent

The best AI agent is Gemini Spark if you want the agent in the cloud and Claude Cowork if you want it on your desktop. These are joint picks rather than a first and a second: Spark runs in a Google Cloud VM, so it keeps working while your laptop is shut, wired into Gmail, Docs and, since early September, Google Photos; Cowork runs on your Mac or Windows machine and drives your local apps and your screen. No independent board rates the two products against each other, so the choice is about where the work has to happen rather than which one scores higher. Price is no longer a tiebreaker either: Spark now reaches the $19.99/month Google AI Pro tier in the US and more than 160 other countries, and Cowork is included at no extra charge on every paid Claude plan from $20/month. ChatGPT Codex Mobile (May 14) is the alternative for coding-agent work, and OpenAI Operator-class browser agents for web tasks. Read the full Gemini Spark vs Claude Cowork comparison.

AgentBest ForWhere It RunsStrengthPrice
Gemini Spark24/7 cloud tasks, Workspace workflowsGoogle Cloud VM (always-on)First true 24/7 agent, deep Workspace integrationFrom $19.99/mo Google AI Pro
Claude CoworkDesktop, app-driving, design + codeYour Mac/Windows desktopDrives local apps, sees your screen$20/mo Claude Pro
ChatGPT Codex MobileCoding agent on phoneOpenAI cloud + iOS/AndroidApprove diffs and redirect work from phoneIncluded in ChatGPT plans
Grok Agentic (Grok 4.6)Real-time research, X scrapingSpaceXAI cloudNative X integration; GDPval-AA v2 Elo 1663$30/mo SuperGrok
OpenAI Operator-classBrowser tasks, web formsOpenAI cloud + your browserWeb automationChatGPT Pro
Runner-up and alternatives: Gemini Spark and Claude Cowork are joint picks split by where the agent runs, not a first and a second. ChatGPT Codex Mobile is the pick for coding agents, and Grok Agentic is the niche pick for real-time research. Google AI Ultra at $99.99 or $199.99 a month now buys Spark higher usage limits rather than access, which starts on the $19.99 Pro tier.
What changed this month

No new consumer agent product shipped, though Gemini Spark did expand: it reached the $19.99/month Google AI Pro tier and gained Google Photos, where it can search, edit, curate and run scheduled workflows. Meta's Muse Code (August 5) is a terminal developer tool rather than a consumer one, so the Spark (cloud) versus Cowork (desktop) choice still drives most agent decisions. The model layer underneath moved a long way. Claude Fable 5.1 (September 1) leads Arena's Agent Arena on task completion at 19.8%, with GPT-6 Astra second at 17.7% and Claude Opus 5 fourth, and it tops Artificial Analysis's AA-Briefcase at 1646 against Opus 5's 1633. On GDPval-AA v2 the two have converged: Fable 5.1 leads at xhigh effort with 1745, but at max effort Opus 5's 1735 is now ahead of Fable 5.1's 1724, and the confidence intervals overlap, so treat that pair as a tie. Two housekeeping notes: Artificial Analysis retired the standalone Agentic Index with Intelligence Index v4.2 on September 4 and now measures agentic work through AA-Briefcase and GDPval-AA v2 instead, and it re-anchored the GDPval-AA v2 Elo scale at the same time, which is why these figures sit well below the ones we quoted in August. Opus 5 is still the one most teams should build on at $5 / $25, half Fable 5.1's price. Meta's Muse Spark 1.3 (September 2) deserves a fairer hearing than its own launch chart gives it: Meta groups six benchmarks under Agent and 1.3 loses all six, but on Artificial Analysis's independent GDPval-AA v2 board it places third among distinct models at 1703, ahead of GLM 5.3, Grok 4.6 and Claude Fable 5, at $1.25 / $4.25. It also wins long context, 98.1 on MRCR's 512K to 1M band against 55.5 for 1.2. Scale SEAL has still rated only the older 1.1, which leads its MCP Atlas tool-use board at 88.1. GLM 5.3 Flash (August 26, MIT) is the open-weight agent model we would host, at AutomationBench 48.8 on Z.ai's own chart where Claude Opus 4.8 scores 41.0; on the Agent Arena's task-completion signal GLM-5.2 now sits seventeenth and Kimi K3 fifth, with DeepSeek V4.1-Flash third overall and the highest open-weight model on that signal. Grok 4.6 remains the cost story rather than a new agent product: a GDPval-AA v2 Elo of 1663, fifth among distinct models, while finishing long-horizon jobs in roughly 53 turns and 0.5 billion input tokens against about 103 turns and 2.0 billion for Opus 5. GPT-6 Astra (September 3) argues for Codex rather than for a new agent product: it leads Artificial Analysis's Terminal-Bench 4.0 and is second on the Agent Arena, but Fable 5.1 keeps AA-Briefcase by 1646 to 1562.

Fello AI running on Mac, iPhone and iPad
All the top AI models, one app

The leading models like ChatGPT, Claude and Gemini, together on Mac, iPhone and iPad.

Free to start, 4.7★ across 27,000+ reviews.

Download Fello AI

Use case guide, September 2026

Best AI for Students

1
GPT-5.6 Luna
Essays & research
2
Gemini 3.6 Flash
STEM + PDFs
3
Claude Sonnet 5
Writing & editing

The best AI for students is GPT-5.6 Luna inside ChatGPT for general coursework and Gemini 3.6 Flash inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Opus 5 as the alternative for essay editing. Most students don't need to pay, and the free tier improved on August 6: GPT-5.6 Luna became the default for ChatGPT Free and Go, with unlimited text chats and a Think button for harder questions, though file uploads and image tools stay capped. Gemini 3.6 Flash is still what the free Gemini app serves, Claude Sonnet 5 is the free Claude default, and DeepSeek V4 is free on DeepSeek's chat site. Luna is the cheapest member of the GPT-5.6 family rather than the strongest, so reach for the Think button or switch to GPT-5.5 when an essay needs more care. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's paid flagship and GPT-6 Astra now reports 97.6% on FrontierMath Tier 4 v2 against GPT-5.5 Pro's verified 39.6%, though Astra is not on ChatGPT's free tier; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).

TaskBest ModelWhyFree?Alternative
Essays & courseworkGPT-5.6 LunaFree default in ChatGPT since August 6; unlimited text chatsYesClaude Sonnet 5 (free Claude)
STEM problem-solvingGPT-5.6 Sol / Qwen 3.7 MaxPaid STEM flagship; Astra reports 97.6% FrontierMath Tier 4 v2 / 97.1 HMMT 2026 FebPro paid / Qwen API paidGemini 3.6 Flash (free)
Research & accuracyGemini 3.1 Pro98% ARC-AGI-1 at $0.52/task, native Google Search groundingYes (Gemini app)Claude Opus 5
Writing editingClaude Sonnet 5Free and default on claude.ai; Claude Fable 5 is the quality leaderYes (Claude free)Claude Fable 5
Multimodal study (PDFs, slides, images)Gemini 3.6 Flash1M context, free in Gemini appYesNotebookLM (Google)
Runner-up and alternatives: Claude Sonnet 5 (free) is the runner-up for essay writing and editing. Gemini 3.6 Flash (free) is the runner-up for multimodal study and PDF ingestion. DeepSeek V4 is the runner-up for problem-solving on a strict zero-cost budget.
What changed this month

The free tier is where this guide moved. On August 6 OpenAI made GPT-5.6 Luna the default for ChatGPT Free and Go and gave both unlimited text chats plus a Think button for harder questions, so the free pick is now Luna rather than GPT-5.5, which stays selectable when an essay needs more care. File uploads, images and other tools are still capped. Nothing moved on the Google side: Gemini 3.8 Flash shipped on September 2 but reaches free users only through AI Studio and the API, so the free Gemini app still serves 3.6 Flash and it keeps the multimodal row. GPT-6 Astra (September 3) is the one to watch for hard problem sets, reporting 97.6% on FrontierMath Tier 4 v2 against GPT-5.5 Pro's verified 39.6%, but it needs a paid ChatGPT plan and no free tier includes it.

Use case guide, September 2026

Best AI for Work & Professionals

1
GPT-5.6
Daily knowledge work
2
Claude Opus 5
Coding & writing
3
Gemini 3.1 Pro
Research & briefings

The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark is the only true 24/7 cloud agent, and since July 30 it reaches the $19.99/month Google AI Pro tier as well as Google AI Ultra.

Use CaseBest ModelKey StatPriceAlternative
Daily knowledge workGPT-5.6ChatGPT's default since July 9; the assistant most people can open$20/mo ChatGPT PlusClaude Opus 5
Coding (proprietary)Claude Opus 5#1 on Arena image-to-WebDev (1,668.6) and #2 on WebDev (1688); Anthropic's recommended default$20/mo Claude ProClaude Fable 5
Coding (cost-effective)Qwen3.8-Max-0902#4 Code Arena: WebDev (1681), behind GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5; API-only on QwenCloud$2 / $6Qwen 3.7 Max; DeepSeek V4.1-Flash
Research & briefingsGemini 3.1 Pro98% ARC-AGI-1 at $0.52/task, Google groundingGoogle AI Pro / UltraClaude Opus 5
Hard math, physics, finance modellingGPT-6 Astra#1 LiveBench Reasoning and ARC-AGI-2 (95%); reports 97.6% FrontierMath Tier 4 v2$100/mo ChatGPT ProGPT-5.6 Sol; Qwen 3.7 Max
Always-on agent workflowsGemini SparkFirst 24/7 cloud agentFrom $19.99/mo Google AI ProClaude Cowork
Live news, X-context creativeGrok 4.6Intelligence Index 44.4 + native X grounding$30/mo SuperGrokGemini 3.1 Pro
All-in-one consolidationFello AIChatGPT + Claude + Gemini + Grok + DeepSeek$9.99/moPay each vendor separately
Runner-up and alternatives: for most professional teams, Claude Opus 5 is the runner-up to GPT-5.6 for daily work and now the coding leader on image-to-WebDev and on price at $5 / $25, with Claude Fable 5 the runner-up for the hardest long-horizon work. Gemini 3.1 Pro is the runner-up for research-heavy roles, and Gemini Spark is the unique pick if you can put a cloud agent to work on long tasks.
What changed this month

Three launches landed in the first three days of September and none of them moves the daily-work pick. Claude Fable 5.1 (September 1) tops Artificial Analysis's Intelligence Index at 53.4 on v4.3 and its AA-Briefcase board at 1646, but Claude Opus 5 stays the recommendation here at $5 / $25, half Fable 5.1's price, and Anthropic's own docs still tell teams to start with Opus 5. Alibaba's Qwen3.8-Max-0902 (September 2) took Arena's WebDev board from Opus 5 by three points and takes the cost-effective coding row at $2 / $6, though it is API-only on QwenCloud with no consumer front-end. GPT-6 Astra (September 3) is the launch that could move the daily-work row, and it has not yet: Astra reached Business and Pro seats on September 4 and Plus a few hours later, but at $10 / $50 against Sol's $4 / $20 it is a costlier default, so GPT-5.6 keeps the row.

Pricing Comparison

AI Model Pricing in September 2026

From $0 free tiers to $199.99/month Google AI Ultra. The most consequential price on this table is Claude Opus 5 at $5 / $25, because it is the third-highest scorer on Artificial Analysis's Intelligence Index at half the list price of the two Fable models around it. Claude Fable 5.1, which took the top of that index on September 1 and still holds it at 53.4 on v4.3, has the same $10 / $50 list price as Fable 5 but cuts cache reads 75% to $0.25, which is where the saving on long agentic runs actually sits rather than on the sticker. Meta's Muse Spark 1.3 lists at $1.25 / $4.25, unchanged across three capability bumps since July, with a Contributor tier at $0.10 / $0.20 for anyone willing to let Meta train on their prompts, and Grok 4.6 lists at $2 / $6, with the caveat that a prompt of 200,000 tokens or more re-bills the entire request at $4 / $12. The cheap end has moved twice. DeepSeek raised both V4 models roughly fourfold on August 16, which cost V4-Flash the price-performance pick, then reversed most of it on September 10: V4-Flash was retired and V4.1-Flash took its place at $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak. The pick stays with GLM 5.3 Flash all the same, now at its plain $0.15 / $0.50 list since the launch promotion expired on September 9, because Artificial Analysis measures it at $0.25 an Intelligence Index task for a score of 41.9 against DeepSeek's $0.27 for 39.5. Off-peak the two are level on input and GLM is cheaper on output; at peak GLM wins both. At the frontier the value pick is Meta's Muse Spark 1.3, at $1.61 an index task for a score of 48.2, ahead of Grok 4.6 on both counts at $1.86 and 44.4. Among closed models GPT-5.6 Luna is the cheapest at $0.20 / $1.20 after OpenAI's July 30 cut, and the cheapest per index task of any model on this page at $0.18. For a deeper breakdown see our full AI Pricing Comparison Guide. The month's new flagship fits that shape rather than changing it: GPT-6 Astra arrived on September 3 at $10 / $50, the same list price as Claude Fable 5.1 and 2.5 times GPT-5.6 Sol, with cached input at $1 and cache writes at $12.50, and Artificial Analysis measures it 64% more expensive per index task than Sol, for 5.7 points more on the index.

ModelInput (per 1M)Output (per 1M)ContextFree Access?
GPT-6 Astra$10.00$50.001,050,000 (922,000 max input)ChatGPT Business/Pro since September 4, Plus hours later; no free tier. API, AWS and Azure still to come
GPT-5.5$5.00$30.001M (400K in Codex)ChatGPT Free; API paid
GPT-5.5 Pro$30.00$180.001MChatGPT Pro from $100/mo ($200 higher-usage tier)
GPT-5.6 Sol$4.00$20.00Not publishedLive in ChatGPT, Codex & API (July 9)
GPT-5.6 Terra$2.00$12.00Not publishedLive in ChatGPT, Codex & API (July 9)
GPT-5.6 Luna$0.20$1.20Not publishedLive in ChatGPT, Codex & API (July 9)
Claude Opus 5$5.00$25.001MClaude Pro/Max default; API paid
Claude Opus 4.8$5.00$25.001MLegacy model at Anthropic; Pro/Max/API
Claude Fable 5.1$10.00$50.001MCache reads $0.25; in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits. Mythos 5.1 bills identically, by invitation only
Claude Fable 5$10.00$50.001MPermanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits
Claude Sonnet 5$2.00$10.001MClaude Free & Pro default; API paid
Claude Sonnet 4.6$3.00$15.001MAPI paid (superseded by Sonnet 5)
Gemini 3.1 Pro$2.00 (≤200K) / $4.00 (>200K)$12.00 (≤200K) / $18.00 (>200K)1MLimited Gemini app; API paid
Gemini 3.8 Flash$0.75 intro / $1.50 from Jan 1, 2027$3.75 intro / $7.50 from Jan 1, 20271MAI Studio, Antigravity, Gemini Enterprise + paid API; in the Gemini app reported for AI Pro/Ultra
Gemini 3.7 Flash$0.75 intro / $1.50 from Jan 1, 2027$3.75 intro / $7.50 from Jan 1, 20271MAI Studio, Android Studio, Antigravity + paid API; in the Gemini app via Spark only (AI Pro/Ultra, excludes EEA/UK/CH/Nigeria)
Gemini 3.6 Flash$0.75 intro / $1.50 from Jan 1, 2027$3.75 intro / $7.50 from Jan 1, 20271MFree Gemini app default; AI Studio; free API tier + paid API
Gemini 3.5 Flash-Lite$0.30$2.501MAI Studio; free API tier + paid API
Qwen3.8-Max-0902$2.00 ($0.17 explicit cache hit, $0.25 implicit)$6.001M (128K output)QwenCloud API only; no consumer chat front-end, no open weights
Qwen 3.7 Max$1.25 promo / $2.50 list$3.75 promo / $7.50 list1M200 free requests/day; API paid beyond that
MiniMax M3$0.30 (50% off $0.60)$1.20 (≤512K)1MOpen weights; hosting costs apply
LongCat-2.0Provider-dependentProvider-dependent1MOpen weights (MIT); hosting costs apply
NVIDIA Nemotron 3 UltraProvider-dependentProvider-dependent1MOpen weights (OpenMDW); hosting costs apply
Qwen 3.5 (open-weight)Self-host / TogetherSelf-host / Together1MOpen weights; hosting costs apply
Nex-N2-ProSelf-host / providersSelf-host / providers1MOpen weights (Apache 2.0); hosting costs apply
Rio 3.5 Open 397BSelf-host / providersSelf-host / providers1MOpen weights (MIT); hosting costs apply
Grok 4.3$1.25$2.501MFree consumer plan; API paid
Grok 4.6$2.00 (<200K) / $4.00 (≥200K)$6.00 (<200K) / $12.00 (≥200K)500KSpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel, Cloudflare
Muse Spark 1.3$1.25 ($0.10 Contributor tier)$4.25 ($0.20 Contributor tier)1MPaid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount
Kimi K3$3.00 ($0.30 cache-hit)$15.001MFree basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License)
Gemini Omni Flash (video)$1.50$17.50 (video output)10-second clipsGemini app / Flow; AI Studio + API
DeepSeek V4-Pro$0.66 off-peak / $1.32 peak ($0.022 cache-hit off-peak)$1.98 off-peak / $3.96 peak1MDeepSeek Chat free; API paid, peak/off-peak tiers since August 16
DeepSeek V4.1-Flash$0.15 off-peak / $0.30 peak ($0.003 cache-hit off-peak)$0.60 off-peak / $1.20 peak1MDeepSeek Chat free; API paid, peak/off-peak tiers; replaced V4-Flash on September 10
Kimi K2.7 CodeProvider-dependentProvider-dependent256KOpen weights; hosting costs apply
GLM-5.2Provider-dependentProvider-dependent1MOpen weights; hosting costs apply
GLM 5.3$1.40 ($0.26 cache-hit)$4.401MZ.ai API, GLM Coding Plan and ZCode; weights on Hugging Face since August 28 under the bespoke GLM 5.3 License
GLM 5.3 Flash$0.15 ($0.03 cache-hit); $0.075 promo$0.50; $0.25 promo1MZ.ai API, GLM Coding Plan, OpenRouter; MIT weights on Hugging Face; promo ends September 9
Tencent Hy4 Preview$0.834 ($0.042 cache-hit)$2.5011M+Open weights (Apache 2.0); Tencent Cloud TokenHub, OpenRouter, WorkBuddy, CodeBuddy
Qwen3.8-Flash-NextProvider-dependentProvider-dependent262K (to 1M)Open weights (Qwen Community License 1.0); hosting costs apply
ERNIE 5.1China-region pricingChina-region pricing256KBaidu free tier
Gemini Spark (agent)Not API-pricedNot API-priced1M (Gemini base)Google AI Pro $19.99, AI Ultra $99.99 or $199.99/mo
Fello AI (aggregator)Routed via appRouted via appModel-dependent$9.99/mo, free tier available

The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. Grok 4.6 works the same way but bites harder: at 200,000 prompt tokens and above, SpaceXAI re-bills the entire request at the higher rate rather than only the overage, so crossing the line by a thousand tokens doubles the cost of the whole call. DeepSeek is the other rate card that needs reading twice: since August 16 both V4 models bill at peak rates from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays, and at exactly half that in every other hour, so the same job can cost twice as much depending on when you run it. If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.

Open-Weight Models

Best Open-Weight Models in September 2026

The best open-weight model in September 2026 is Kimi K3, and the case for it has narrowed. It is still the highest-placed open model on Arena, fifth on WebDev at 1674 and fifth on the Agent Arena's task-completion signal, but it no longer holds the highest open Intelligence Index: on Artificial Analysis's v4.3 index GLM-5.3 scores 44.9 against Kimi K3's 43.8, and GLM-5.3 also leads it on AA-Briefcase, 1504 to 1488, and on Terminal-Bench 4.0 by a wide margin, 42% to 13%. No open model sweeps, so we keep Kimi K3 on its Arena placement and say plainly that the measured boards now point at GLM-5.3, whose licence is a bespoke Z.ai document rather than MIT. One catch decides which you should actually use. The K3 download is 96 safetensors shards and about 1.56 TB, which needs a multi-node GPU cluster rather than a workstation. So the practical recommendation for teams running their own weights is still GLM 5.3 Flash (Z.ai, MIT), released on August 26 and scoring 41.9 on v4.3, almost eight points above GLM-5.2 at less than half the size, 320B total with 18B active. It is the first natively multimodal model in the GLM-5 line, it takes a 1M-token context, and the weights were on Hugging Face under MIT the day it launched. The cheap open tier changed again on September 10: DeepSeek retired V4-Flash 0731 and replaced it with DeepSeek-V4.1-Flash, 552B with 8B active at prefill and 16B at decode, natively multimodal, MIT on the day, scoring 39.5 on v4.3 against 34.5 for the build it replaces, and priced back down to $0.15 / $0.60 off-peak. It is also third on the Agent Arena's task-completion signal, the highest-placed open model there. Two of the three releases that closed August now have independent ratings: GLM 5.3 published its weights on August 28 under a bespoke licence with a clause requiring any model-as-a-service operator above $10 billion in revenue to pass a Z.ai security review first, and Artificial Analysis scores it 44.9; Alibaba's Qwen3.8-Flash-Next, a 125B mixture-of-experts with only 6B active that previews the Qwen4 architecture, scores 39.9 and sits ninth on Arena's WebDev board. Tencent's Hy4 Preview, 770B total and 49B active under Apache 2.0 and still the largest permissively licensed model published, has no Artificial Analysis rating but now sits eleventh on Arena WebDev at 1624. Note also that GLM 5.3 Flash is not a trimmed GLM 5.3 but a separate model on a newly trained base, which is why it could ship under plain MIT while the flagship could not.

ModelBest ForKey BenchmarkContext / LicenseWhere To Run
Kimi K3Highest-scoring open model, agentic and web workII 43.8 (v4.3); #5 Arena WebDev, the best open placement there; #5 Agent Arena; 2.8T/104B active1M / Kimi K3 LicenseHugging Face (96 shards, 1.56 TB), Moonshot API, providers
GLM 5.3 FlashBest open model most teams can actually runII 41.9 (v4.3), on the intelligence-versus-cost Pareto frontier at about $0.25 an index task; 320B/18B active, natively multimodal1M / MITHugging Face, Z.ai API ($0.15/$0.50), OpenRouter
GLM-5.2Fallback with an independent track recordII 34.0 (v4.3); #19 Arena WebDev (1,591.8), #17 Agent Arena; 744B/40B active1M / MITZ.ai, Hugging Face, OpenRouter
GLM 5.3Open since August 28, but on a bespoke licenceII 44.9 (v4.3), the highest open score we found; same 744B base as GLM-5.2, post-trained only, published checkpoint counted at 753B; CyberGym 84.5% and Terminal-Bench 3.0 28.3 (vendor figures)1M / GLM 5.3 License (MaaS above $10B revenue needs a Z.ai security review)Hugging Face, Z.ai API ($1.40/$4.40), GLM Coding Plan, ZCode
DeepSeek V4.1-FlashBest open value if you host it yourselfII 39.5 (v4.3) against 34.5 for the V4-Flash 0731 build it replaces; 552B, 8B active at prefill and 16B at decode1M / MITDeepSeek API ($0.15/$0.60 off-peak, $0.30/$1.20 peak since September 10), local
Tencent Hy4 PreviewLargest permissively licensed model yet770B/49B active; 2.99/4.00 on Tencent's own 203-task expert panel vs Kimi K3 2.94 and GLM 5.3 2.92 (vendor); no Artificial Analysis rating; #11 Arena WebDev (1624)1M+ / Apache 2.0Hugging Face (BF16 and FP8), Tencent Cloud TokenHub, OpenRouter
Qwen3.8-Flash-NextPreview of the Qwen4 architecture125B total plus a 51B N-gram embedding, 6B active; 91.7 GPQA Diamond and 62.5 SWE-Bench Pro (vendor); II 39.9 (v4.3); #9 Arena WebDev (1635)262K to 1M / Qwen Community License 1.0Hugging Face, providers, self-host
LongCat-2.0Frontier open coder trained on Chinese chipsII 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active1M / MITHugging Face, GitHub, OpenRouter
MiniMax M3Cheap frontier-class multimodalII 44 (v4.1), 59% SWE-Bench Pro, multimodal1M / license TBDHugging Face, API $0.30/1M (50% off)
Nex-N2-ProStrongest open coding scoreII 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B activeQwen-based / Apache 2.0Hugging Face, providers, self-host
Kimi K2.7 CodeStrongest commercially-licensed open coder+21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active256K / Modified MITHugging Face, DeepInfra, providers
DeepSeek V4-ProAgentic real-world workII 44 (v4.1), 1.6T/49B active1M / MITDeepSeek API ($0.66/$1.98 off-peak, $1.32/$3.96 peak since August 16), local
Hy3Newest permissive-licence entrantII 41 (v4.1); #22 on Arena WebDev (1,516.4)Apache 2.0Hugging Face, providers, self-host
InklingThinking Machines' first open modelII 41 (v4.1), agentic 32.3, released July 15Open weightsHugging Face, providers, self-host
Inkling SmallSame family at a quarter the sizeII 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio inApache 2.0Hugging Face (BF16 and NVFP4), providers, self-host
NVIDIA Nemotron 3 UltraNVIDIA-tuned, fully permissive licenseII 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active1M / OpenMDWOpenRouter, Hugging Face, AWS (8x B200 self-host)
Qwen 3.5 (397B / 17B active)Multimodal, fast decode88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v61M / openTogether, OpenRouter, local
Qwen3.6-35B-A3BEfficient open agentic coder (3B active)86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active262K (to 1M YaRN) / Apache 2.0Hugging Face, OpenRouter, local
Qwen3.6-27BLaptop-runnable dense coder87.8 GPQA Diamond, dense 27B, multimodal256K / Apache 2.0Local Mac/PC, Hugging Face, OpenRouter
Rio 3.5 Open 397BQwen 3.5 fine-tune, multilingual reasoning70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5397B/17B active / MITHugging Face, providers, self-host
Llama 4 MaverickMeta-line flagship17B active / 400B total paramsLlama 4 licenseMeta cloud, Hugging Face, local
NVIDIA Nemotron 3 Nano OmniEdge / low-powerMultimodal, very small footprintCompact / openLocal, NVIDIA tool

Licensing matters as much as raw score here, and August widened the gap between the two groups. Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM 5.3 Flash (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Hy4 Preview (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use with no revenue test. The rest carry conditions worth reading before you build on them: MiniMax M3 and MiniMax H3 ship under their own community licences, H3 additionally requiring an application from the US, EU, UK and South Korea; Kimi K3 sits on a custom licence with a revenue threshold for anyone reselling it as a service; GLM 5.3 requires a Z.ai security review of any model-as-a-service operator above $10 billion in revenue; and the Qwen Community License 1.0 on Qwen3.8-Flash-Next requires a separate licence from Qwen to run a model-as-a-service or an AI work-assistant business, though internal use is exempt. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.

How We Evaluate

Benchmarks, Prices, and Hands-On Use

Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 2.1 board, covering the Intelligence and Agentic indexes, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.

Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We rank on measured scores: a category crown moves as soon as a model out-scores the holder on the boards that define that category, without waiting for the vote-based boards to catch up, and a model that is not yet widely available is flagged on its card rather than denied one. We re-check ranks as well as figures, because a score can stay correct while the board around it moves. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.

Fello AI running on Mac, iPhone and iPad
Not sure which model to pick?

Fello AI brings the top AI models together in one lightweight native app.

Switch between them anytime, no separate subscriptions.

Download Fello AI
Frequently Asked Questions

Common Questions

What is the best AI model right now in September 2026?
It depends on the task. On overall benchmark score, Claude Fable 5.1 (September 1) is #1 on Artificial Analysis's Intelligence Index at 53.4 on v4.3 and #1 on AA-Briefcase at 1646, with GPT-6 Astra second on the index at 52.8 and Claude Opus 5 third at 50.7. Arena has now rated Fable 5.1: second on both coding boards and fifth on text. Meta's Muse Spark 1.3 is the value pick at the frontier, at $1.61 an index task for a score of 48.2, ahead of Grok 4.6 on both counts. For daily chat, GPT-5.6 has been ChatGPT's default since July 9 and is the assistant most people can actually open, although Claude Fable 5 leads Arena's text leaderboard. For coding, GPT-6 Astra leads Artificial Analysis's Terminal-Bench 4.0 at 60% and both of Arena's coding boards, with Claude Opus 5 the value pick at half the list price. For writing, Claude Fable 5 is #1 on Arena's text leaderboard, with Fable 5.1 ahead of it on Humanity's Last Exam (59%) and on AA-Omniscience accuracy (67%). For accuracy, Gemini 3.1 Pro ties the human panel on ARC-AGI-1 at 98% for $0.52 a task. For hard math and reasoning, GPT-6 Astra leads LiveBench Reasoning and ARC-AGI-2, with Claude Fable 5.1 top of LiveBench Mathematics. For images, ChatGPT Images 2.5 leads both of Arena's boards, and for video Gemini Omni 1.1 Flash leads text-to-video. For agents, Gemini Spark is the 24/7 cloud agent and Claude Cowork the desktop one.
Is GPT-6 out, and is it the best model now?
GPT-6 is out. OpenAI launched it as GPT-6 Astra on September 3, 2026, which settles the year-long question of whether Astra would ship as GPT-6 or as another GPT-5 point release. It is not the top model on every independent board. Artificial Analysis scores it 52.8 on Intelligence Index v4.3, 5.7 points clear of GPT-5.6 Sol, 0.6 behind Claude Fable 5.1 at 53.4 and 2.1 ahead of Claude Opus 5 at 50.7, while charging $10 / $50 per 1M tokens against Sol's $4 / $20. It is stronger on coding: Artificial Analysis has since retired that Coding Agent Index, and on its own Terminal-Bench 4.0 run Astra leads outright at 60% against Fable 5.1's 55%, on roughly a third of Sol's output tokens. It also roughly halves hallucination on AA-Omniscience, from 92% to 51% at max effort. You need a paid plan: Astra is the first model OpenAI has designated Critical for cybersecurity, so approved defenders in its Daybreak program got it first, with ChatGPT Business and Pro following on September 4 and Plus a few hours later. The API, AWS and Azure are still rolling out, and nothing has been announced for the free tier. Our full GPT-6 Astra breakdown covers the benchmarks and the caveats.
What is Claude Opus 5?
Claude Opus 5 is Anthropic's July 24, 2026 flagship, and it was the #1 model on Artificial Analysis until Claude Fable 5.1 arrived on September 1. It now sits third on the Intelligence Index at 50.7, against Claude Fable 5.1's 53.4 and GPT-6 Astra's 52.8, second on AA-Briefcase at 1633 against 1646, and at max effort it actually leads the GDPval-AA v2 professional-deliverables board with 1735 against Fable 5.1's 1724, on overlapping confidence intervals, well clear of Grok 4.6's 1663 and Claude Fable 5's 1632. API pricing is $5 / $25 per 1M tokens, the same as Opus 4.8 and half the cost of Fable 5, with a 1M-token context window and a May 2026 knowledge cutoff. ARC Prize independently confirms Anthropic's launch claim on ARC-AGI-3, where Opus 5 scores 30% against 8% for the next-best model, roughly 3.75x. Arena now places it third on both coding boards, behind GPT-6 Astra and Claude Fable 5.1, first on Document, fourth on the Agent Arena's task-completion signal and tenth on text. One thing to keep in mind: Artificial Analysis measures its hallucination rate at 50%, which is why it holds no accuracy crown on this page.
What is Claude Fable 5.1?
Claude Fable 5.1 is Anthropic's September 1, 2026 release and the highest-scoring model Artificial Analysis has measured, at 53.4 on its Intelligence Index v4.3 at max effort and first on AA-Briefcase at 1646. It shipped alongside Claude Mythos 5.1, which is the same model with the biology and cybersecurity safeguards relaxed, available by invitation only and with no published eligibility criteria. Pricing is $10 / $50 per 1M tokens, unchanged from Fable 5, but cache reads fall 75% to $0.25, on a 1M-token context window with a June 2026 knowledge cutoff. It is included in Claude Max and Team Premium at roughly 50% of weekly limits and reachable from Pro through usage credits. Anthropic still recommends starting with Claude Opus 5 for most workloads, since it costs half as much and runs faster. Our full Claude Fable 5.1 and Mythos 5.1 breakdown covers the system card and the safeguard split.
Is Claude Fable 5 back?
Yes. Anthropic redeployed Claude Fable 5 on July 1, 2026 after the US government lifted the export-control restriction it had imposed on June 12. It is available again on the Claude API, Claude.ai, Claude Code, and Claude Cowork. Anthropic made Fable 5 permanent in the paid plans on July 20, 2026. Max and Team Premium include it at roughly 50% of regular usage limits; Pro and Team Standard reach it through usage credits with a one-time $100 starting credit. API pricing is $10 / $50 per million tokens. Fable 5 is a Mythos-class model built for long-horizon agentic work with a 1M-token context, and it is the runner-up for coding behind Claude Opus 5, at #2 on Arena's image-to-WebDev board (1,625.7) and #4 on WebDev.
What is GPT-5.6 and can I use it?
Yes, as of July 9, 2026. GPT-5.6 is OpenAI's next-generation model family, and after a two-week gated preview that began June 26 behind a US-government safety review, it reached general availability on July 9. There are three tiers, least to most capable: Luna, the fast, cheapest tier ($0.20 / $1.20 per 1M tokens); Terra, a balanced everyday model OpenAI says matches GPT-5.5 ($2 / $12); and Sol, the flagship tuned for biology, chemistry, and cybersecurity ($4 / $20). OpenAI cut Luna by 80% and Terra by 20% on July 30, 2026 and left Sol alone, then cut Sol by more than 20% on August 21, 2026 as a promotional rate for about three months. No ChatGPT subscription price has changed. It is now live across ChatGPT, Codex, and the API as OpenAI's default model. One caveat: OpenAI's system card and the external evaluator METR flagged elevated "scheming" behaviour in Sol, so treat it carefully for high-stakes factual work until independent results settle.
Is Grok 4.6 out yet?
Yes. SpaceXAI released Grok 4.6 on August 12, 2026, replacing Grok 4.5 as the flagship just over a month after it shipped. It is not a new foundation model: the company kept the Grok 4.5 base and improved it through additional post-training, including reinforcement learning in agentic environments, and published no new architecture or parameter count. Artificial Analysis scores it 44.4 on Intelligence Index v4.3, up five points from Grok 4.5's 39.1, behind GPT-5.6 Sol at 47.1, Claude Fable 5.1 at 53.4, GPT-6 Astra at 52.8, Claude Opus 5 at 50.7 and Claude Fable 5 at 49.7. Pricing is unchanged at $2 / $6 per 1M tokens on a 500K-token context window, though prompts of 200,000 tokens and above re-bill the whole request at $4 / $12. It is live in the SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare. Grok 4.6 is also our creativity pick, on product grounds rather than prose quality; for the best-written output, Claude Fable 5 wins outright. Full numbers are in our Grok 4.6 breakdown.
Which AI is the best for coding?
GPT-6 Astra is the best for coding. It leads Artificial Analysis's own Terminal-Bench 4.0 at 60% against Claude Fable 5.1's 55%, and it tops both of Arena's coding boards, WebDev at 1800 and image-to-WebDev at 1733. Claude Fable 5.1 is the runner-up and keeps the composite measures, the Intelligence Index at 53.4 and AA-Briefcase at 1646. Claude Opus 5 is the value pick at $5 / $25 per 1M tokens, half the list price of either, with 49% on Terminal-Bench 4.0 and third place on both Arena coding boards, and Anthropic's own docs still say to start with Opus 5 for complex agentic coding. Note that Artificial Analysis has retired both its Coding Index and its Coding Agent Index, so the composite coding rankings this page used to quote no longer exist. Alibaba's Qwen3.8-Max-0902 held Arena's WebDev board for about three days in early September and is now fourth. GLM 5.3 Flash (MIT, August 26) is the price-performance pick at $0.15 / $0.50 and the cheapest serious coder you can host, at 33% on Terminal-Bench 4.0 for two cents a task, though GLM-5.3 reaches 42% on the same harness.
Which AI is the best for writing?
Claude Fable 5.1 is the best for writing, topping Humanity's Last Exam at 59.1% and the GDPval-AA v2 professional-deliverables board. Claude Fable 5 is the pick if you weight human votes: it still leads Arena creative writing and LiveBench Language (90.7), though it now sits #8 on EQ-Bench Creative Writing v3. It costs $10 / $50 per 1M tokens. Claude Sonnet 5 is the value pick and what most people should actually use, since it is free and default on claude.ai at $2 / $10, a launch rate Anthropic made permanent in August 2026, though it ranks #53 on Arena creative writing. Kimi K3 is fourth on EQ-Bench at 2070.6 if you want distinctive fiction, Claude Fable 5.1 tops the GDPval-AA v2 professional-deliverables board at 1853 with Claude Opus 5 second at 1824, GPT-5.5 is the alternative for fact-anchored business writing, and Gemini 3.8 Flash is the price-performance pick for bulk content.
What is the best open-weight AI model in 2026?
Kimi K3 is still our open-weight pick, but the case has narrowed. It is the highest-placed open model on Arena, fifth on WebDev at 1674 and fifth on the Agent Arena's task-completion signal, yet it no longer holds the highest open Intelligence Index: GLM-5.3 scores 44.9 on Artificial Analysis's v4.3 index against Kimi K3's 43.8, and leads it on AA-Briefcase, 1504 to 1488, and on Terminal-Bench 4.0, 42% to 13%. No open model wins everything, so we keep Kimi K3 on its Arena placement and say plainly that the measured boards point at GLM-5.3, whose licence is a bespoke Z.ai document rather than MIT. The practical catch is size: K3 is 96 shards and about 1.56 TB under a custom Kimi K3 License, so GLM 5.3 Flash (August 26, MIT) is the model most teams can actually host, at Intelligence Index 41.9 with 320B total and 18B active parameters. The cheap tier changed on September 10, when DeepSeek retired V4-Flash 0731 and replaced it with DeepSeek-V4.1-Flash: 552B with 8B active at prefill and 16B at decode, MIT, natively multimodal, 39.5 on v4.3 against 34.5 for the build it replaces, and back down to $0.15 / $0.60 off-peak. Two of August's closing releases now have ratings as well: Qwen3.8-Flash-Next at 39.9, and Tencent's Hy4 Preview, still unrated by Artificial Analysis but eleventh on Arena's WebDev board. Worth knowing honestly, no open model is first on any Arena board we track.
What is the cheapest frontier-class AI model?
On API pricing per million tokens, GPT-5.6 Luna is the cheapest closed frontier-class model at $0.20 / $1.20, after OpenAI cut it 80% on July 30, 2026, and Artificial Analysis scores it 37.5 on its v4.3 index against Gemini 3.6 Flash's 34.3. Cheaper still is GLM 5.3 Flash, our price-performance pick since August, now at its plain $0.15 / $0.50 list since the launch promotion expired on September 9, scoring 41.9 under a plain MIT licence you can self-host. DeepSeek held that pick until it raised rates roughly fourfold on August 16, and it came close to taking it back on September 10, when V4.1-Flash replaced V4-Flash at $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak: off-peak the two now tie on input, with GLM cheaper on output. Artificial Analysis still separates them on cost per index task, $0.25 for GLM against $0.27 for DeepSeek. Google's cheap tier improved again on September 2, with Gemini 3.8 Flash scoring 41.2 at the same introductory $0.75 / $3.75, though that rate doubles on January 1, 2027. MiniMax M3 lists at $0.30 per million input tokens on a permanent 50% discount, with LongCat-2.0 as a strong MIT alternative. Measured per Intelligence Index task rather than per token, the cheapest model on this page is GPT-5.6 Luna at $0.18.
Which AI models are free?
ChatGPT Free now defaults to GPT-5.6 (with GPT-5.5 still available) under usage limits. Gemini Free runs Gemini 3.6 Flash in the Gemini app and Google AI Studio. Claude Free runs Claude Sonnet 5 (the new default) with daily limits. DeepSeek Chat runs DeepSeek V4 free on the DeepSeek website. Grok has a limited free consumer plan (X Premium is a paid add-on). Qwen 3.5, Qwen3.8-Flash-Next, NVIDIA Nemotron 3 Ultra, MiniMax M3, LongCat-2.0, Kimi K2.6, DeepSeek V4, GLM-5.2, GLM 5.3, GLM 5.3 Flash and Tencent's Hy4 Preview are open-weight and free to self-host, though the licences differ and only some of them are plain MIT or Apache. Qwen 3.7 Max is API-only with no consumer chat front-end, but Alibaba now includes 200 free model requests per day. Kimi has a free basic tier in its app, with heavier agentic use metered.
What is Gemini Spark and which plan do you need?
Gemini Spark is Google's first 24/7 cloud-resident AI agent, launched at Google I/O on May 19, 2026 and no longer an Ultra exclusive: since July 30, 2026 it also reaches the $19.99/month Google AI Pro tier, in the US and more than 160 further countries, while Google AI Ultra at $99.99 or $199.99/month carries the wider global rollout. Google AI Plus, the free tier and work or school accounts do not include it. Spark is built on Gemini base models with Google's Antigravity harness on a Google Cloud VM, integrates with Gmail, Google Docs, and other Google Workspace apps, and can interact with Chrome and Android's Halo system on the device side. It is worth the spend for users who have repeatable long-running workflows (inbox triage, research roll-ups, scheduled tasks). For one-off tasks, Claude Cowork at $20/month covers most desktop-agent needs.
What is Fello AI?
Fello AI is an AI chatbot for Mac, iPhone, and iPad that lets you use all top AI models like ChatGPT, Claude, Gemini, Grok, and DeepSeek in one app, with models updated regularly so you always have the latest. It is $9.99/month with a 4.7-star rating across 27,000+ reviews.
How often do you update this page?
We update this page at least monthly and within 24-48 hours of any major model launch.
Related Articles

Try every model.
One beautiful app.

Every model from this guide in one native app: ChatGPT, Claude, Gemini, Grok, and DeepSeek. Free to start.

Fello AI running on Mac, iPad, and iPhone

4.7 rating·27,000+ reviews·Free to start