Updated October 7, 2026

Best AI Models in 2026

Rankings, comparisons, and deep dives, updated monthly as new models ship.

Google announced Gemini 4 Argon on September 30, its first flagship since Gemini 3.1 Pro, and the independent boards put it straight into the top tier: Artificial Analysis scores it 52.6 on Intelligence Index v4.3.2, level with GPT-6 Astra and behind only Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1, with the lowest hallucination rate it has measured among leading models, and Arena ranks it first on its text and creative-writing boards. It wins no category on this page, for two reasons: most of the boards our categories are decided on have not rated it yet, and only vetted cyber defenders in Google's Fairwind Program can use it. Google says paid API customers and Google AI Ultra subscribers come next, with no date, at an introductory $2 / $10 per 1M tokens.

Anthropic still holds most of the measured crowns. Claude Opus 5.5 (September 22) is #1 on the Intelligence Index at 57.6 and on Arena's WebDev board, and keeps the coding crown at $4 / $20; Claude Sonnet 5.5 (September 28) is second at 56.0, leads Terminal-Bench 4.0 at the same $2 / $10 as Sonnet 5, and is the Claude model you can use free on claude.ai. Claude Fable 5.1 keeps writing and accuracy on the boards those categories are decided on. This page ranks on measured scores, so a crown moves as soon as a model out-scores the holder, without waiting for the vote-based boards to catch up. Our guide to AI benchmarks explains how that index is built and why its version number matters.

OpenAI's answer is price. GPT-6 Sol and GPT-6 Luna (September 22) halved its API rates, and GPT-6.1 Sol (September 29) scores 51.8 at the same $2 / $10, sixth among distinct models but the cheapest in the top tier at $0.72 an index task. ARC Prize now puts it second on ARC-AGI-2 at 94.2%, 0.8 points behind GPT-6 Astra, and Arena fourth on WebDev. None of the GPT-6 models is in ChatGPT's chat window yet: they run in ChatGPT Work and Codex.

SpaceXAI's Grok 4.7 (September 21) reached the API, Cursor and Grok Build at an unchanged $2 / $6 while the Grok app still serves 4.6, and Xiaomi's MiMo-V2.6-Pro (September 21, plain MIT) is still the highest-scoring open model at 46.3 and $0.13 an index task. On October 6, Mistral released Mistral Large 4, the highest-scoring model from the US or Europe that promises open weights, and Google replaced Nano Banana 2 with Nano Banana 2.1. Neither takes a crown, but video changes hands: Artificial Analysis has now rated Gemini Omni 1.1 Flash, eighth on its text-to-video board, and MiniMax H3 takes the category. Below are the category winners for October 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks and pricing are updated within 48 hours of any major model launch.

October 2026 Category Winners
Writing
Claude Fable 5.1
#2 EQ-Bench Creative Writing and #3 LiveBench Language
In the top ten of all three prose boards, second on EQ-Bench and third on LiveBench Language. Best value: Claude Sonnet 5.5, free on claude.ai.
Deep dive →
Chat & Daily Assistant
GPT-5.6
ChatGPT default since July 9
The best assistant most people can actually open, balancing capability, speed, and reach.
Deep dive →
Images
ChatGPT Images 2.5
#1 and #2 on all four image boards we track
Its two models hold the top two places on every image board we track, Arena's and Artificial Analysis's alike, and it is the default on every ChatGPT tier including free. Best value: Flare, the faster of the two, at the same token price as GPT Image 2.
Deep dive →
Video
MiniMax H3
#1 Arena image-to-video (1495), top four on three of four boards
2K clips of up to 15 seconds with native audio and open weights. Best for long takes: Gemini Omni 1.1 Flash, still #1 on Arena text-to-video.
Deep dive →
Coding
Claude Opus 5.5
#1 Arena WebDev (1815) and #1 Intelligence Index (57.6)
Anthropic's September 22 flagship now leads Arena's WebDev board, beats Sonnet 5.5 by 15 points on LiveBench Agentic Coding and costs less per index task than Sonnet 5.5 at max effort. Best value: Claude Sonnet 5.5 at $2 / $10, which leads Terminal-Bench 4.0 (63.6%) and LiveBench Coding.
Deep dive →
Creativity
Grok 4.7
Fewest content restrictions + native real-time X
The pick for edgy, on-trend work: fewest guardrails and native X integration. Grok 4.7 is on the API, Cursor and Grok Build; the $30/month Grok app still serves 4.6.
Deep dive →
Accuracy & Research
Claude Fable 5.1
Highest factual accuracy Artificial Analysis has measured (67%)
Tops the boards that measure how often a model is simply wrong. Best value: GPT-6.1 Sol, 62% factual accuracy at $0.72 an index task.
Deep dive →
Problem Solving
GPT-6 Astra
#1 LiveBench Reasoning (92.65) and #1 ARC-AGI-2 (95%)
Took the hardest reasoning boards off GPT-5.6 Sol days after launching. Best value: GPT-6.1 Sol at $2 / $10, 0.02 behind it on LiveBench Reasoning and second on ARC-AGI-2.
Deep dive →
AI Agents
Gemini Spark
Always-on cloud agent from $19.99/month
Runs continuously in a Google Cloud VM, deeply integrated with Gmail, Docs, and Chrome.
Deep dive →

Want the top AI models without juggling separate subscriptions? Fello AI brings the leading models together in one native app for Mac, iPhone and iPad.

Download Fello AI
What's New in October 2026
Oct 6 Mistral Mistral Large 4 is the strongest model from the US or Europe that promises open weights, but you cannot download it yet New
Mistral released Mistral Large 4, nicknamed Le Chonk, as a public preview API on October 6, 2026. It is a 1.05-trillion-parameter mixture of experts with 49 billion active at launch (Mistral's docs card now says 52 billion), reads text and images, and lists at $1.36 / $4.18 per 1M tokens, though Mistral's own model page shows half that without saying why or for how long. Artificial Analysis scores it 38.4 on Intelligence Index v4.3.2, far above the best US open model, NVIDIA's Nemotron 3 Ultra at 23, but five to eight points behind GLM-5.3, Kimi K3 and MiMo-V2.6-Pro, and it is verbose enough that an index task costs $1.13. Arena has not rated it yet. Mistral says the weights follow by the end of October and has not named a licence. Our Mistral Large 4 breakdown covers the benchmarks and the cyber-defense pitch.
Oct 6 Google Nano Banana 2.1 replaces Nano Banana 2 at half the price per image and lands fifth on Arena text-to-image New
Google released Nano Banana 2.1 on October 6, 2026 as the direct successor to Nano Banana 2 (Gemini 3.1 Flash Image) and deprecated Nano Banana 2 the same day, with an API shutdown on October 29. In the Gemini API an image costs $0.0336 per 1K image against $0.067 before, but input tokens triple to $1.50 per 1M and the cheap 512px size is gone. Arena ranked it fifth on text-to-image at 1328.1, in the same rank band as MAI-Image-2.6 and Grok Imagine Image 2.0, and sixth on image editing at 1428.1, the highest of any Google model on both boards, while ChatGPT Images 2.5 still holds the top two places on each. Artificial Analysis has not rated it yet. It is in the Gemini app, AI Mode in Google Search, Google AI Studio, Flow and the Gemini API. Our Nano Banana 2.1 guide covers what changed and the real cost.
Sep 30 Google Gemini 4 Argon tops Arena's text board and matches GPT-6 Astra on the Intelligence Index, but only cyber defenders can use it New
Google announced Gemini 4 Argon on September 30, 2026, the first model of the Gemini 4 generation and its first flagship since Gemini 3.1 Pro in February. It is a reasoning model aimed at long, complex work, coding, legal and financial research and cybersecurity defense, and it can write up to 1M output tokens in one response, up from 64K. Artificial Analysis scores it 52.6 on Intelligence Index v4.3.2 at high effort, level with GPT-6 Astra at 52.7 and fifth among distinct models behind Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1, at $1.99 an index task on the introductory price. Its standout result there is reliability: a 15% hallucination rate on AA-Omniscience, the lowest Artificial Analysis has measured among leading models, though it answers only 50% of the questions correctly against 67% for Fable 5.1. It also leads AutomationBench-AA at 77.5%, ahead of Claude Sonnet 5.5's 71.3%, but trails on Terminal-Bench 4.0 at 57.1% against Sonnet 5.5's 63.6%. Within a day Arena ranked it first on text at 1524.8 and on creative writing and math, but only eighth on WebDev at 1679. On Google's own table it leads DeepSWE v1.1 at 77.9% and the Vals Index, and loses Terminal-Bench 4.0 to Claude Opus 5.5 and FrontierSWE v2 to GPT-6 Astra. Access is the catch: Argon went only to vetted cyber defenders in Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers next and no date given. The introductory price is $2 / $10 per 1M tokens, with cached input 95% off, against a standard $4 / $20. No crown on this page moves until people can use it and the other boards rate it. Full detail in our Gemini 4 Argon breakdown.
Sep 29 OpenAI GPT-6.1 Sol reaches fifth on the Intelligence Index at $2 / $10, the cheapest per task in the top tier New
OpenAI released GPT-6.1 Sol at DevDay on September 29, 2026, a week after GPT-6 Sol, at the same $2 / $10 per 1M tokens with cached input halved to $0.10. OpenAI says it nearly matches GPT-6 Astra on agentic coding, computer use and professional work at a fifth of Astra's token prices, and the independent numbers put it close rather than level. Artificial Analysis scores it 51.8 on Intelligence Index v4.3.2 at max effort, fifth among distinct models behind Claude Opus 5.5 at 57.6, Claude Sonnet 5.5 at 56.0, Claude Fable 5.1 at 53.4 and Astra at 52.7, against 47.5 for GPT-6 Sol, and an index task costs $0.72, against $3.26 for Astra and $5.98 for Opus 5.5. LiveBench puts it second on Reasoning at 92.63, 0.02 behind Astra, third on Mathematics at 96.83 and second on Language at 90.13, but well outside the top ten on Coding at 80.7. It takes 56.1% on Terminal-Bench 4.0 against GPT-6 Sol's 43.9% and scores 1575 on GDPval-AA v2.1, above Astra's 1542. Arena, EQ-Bench and ARC Prize have not rated it yet. It is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, not in Chat, and every GPT-6 tier bills 2x input and 1.5x output on the whole request once a prompt passes 272K tokens. OpenAI shipped no GPT-6.1 Astra beside it; the Wall Street Journal reported that its October release was cancelled after internal safety tests. The same event launched OpenAI dots, always-on agents that run on GPT-6 Astra, and a $500 ChatGPT Pro plan. Since then Arena has put it third on WebDev at 1759, ARC Prize has scored it 94.2% on ARC-AGI-2, second only to Astra, and Gemini 4 Argon has pushed it to sixth on the index.
Sep 28 Anthropic Claude Sonnet 5.5 lands second on the Intelligence Index at 56.0, at Sonnet's $2 / $10 New
Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 family, at the same $2 / $10 per 1M tokens as Sonnet 5, with cache reads at $0.20. Anthropic says it runs more than 30% faster and costs up to 30% less per task, and the independent numbers back the capability claim: Artificial Analysis measures it at 56.0 on Intelligence Index v4.3.2 at max effort, second only to Claude Opus 5.5 at 57.6 and ahead of Claude Fable 5.1 at 53.4 and GPT-6 Astra at 52.7, against 38.2 for Sonnet 5. It takes Artificial Analysis's Terminal-Bench 4.0 run at 63.6% against Opus 5.5's 59.6%, LiveBench Coding at 91.4 against 89.3 and AutomationBench-AA at 71.3% against 69.5%, and it sits two Elo points behind Opus 5.5 on GDPval-AA v2.1, 1844 to 1846. The catch is effort. At max, Sonnet 5.5 spent 410 million tokens on the index and a task costs $7.60, against $5.98 for Opus 5.5 and $5.09 for Sonnet 5. At xhigh it scores 51.9 for $2.74 a task, close to GPT-6 Astra's 52.7 for less than Astra's $3.26, and at high 46.7 for $1.08. The Claude apps run it at Medium effort by default and the API at High. It is weaker where Opus 5.5 is strongest: 56.3 on LiveBench Agentic Coding against 71.7, and 32.3 on the hallucination-penalised AA-Omniscience index against 46.4. Arena has not ranked it yet. Anthropic says anyone can chat with Sonnet 5.5 on claude.ai, and higher-risk cybersecurity requests fall back to Sonnet 5, which is now a legacy model. Arena has since ranked it fifth on WebDev, at 1709.
Sep 22 OpenAI GPT-6 Sol and GPT-6 Luna halve OpenAI's API prices, and land everywhere except ChatGPT's chat window New
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, about ninety minutes after Anthropic shipped Claude Opus 5.5, and the launch is an argument about cost per task rather than about the top of the board. Sol drops from $4 / $20 to $2 / $10 per 1M tokens and Luna from $0.20 / $1.20 to $0.10 / $0.50, with cached input reads discounted 90%, to $0.20 and $0.01. OpenAI labels both cuts 50%, but Luna's output falls 58.3%, and the comparison is against GPT-5.6 promotional pricing rather than list. Two spokespeople and OpenAI's Tibo Sottiaux all say the new rates are permanent. Where you can use them is the part every summary flattened: the models are in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, Free and Go users reach Luna in the desktop app, and OpenAI states plainly that they are "not yet available in Chat". Independently, Artificial Analysis scores Sol 47.5 on Intelligence Index v4.3.2 at max effort against GPT-5.6 Sol's 47.0, at $1.06 an index task against $1.99, with Terminal-Bench 4.0 at 43.9% against 39.9% and the hallucination rate down from 92% to 60%. Luna is level with GPT-5.6 Luna at 37.3 but runs an index task for $0.07 against $0.18, which makes it the cheapest model per index task on this page. ARC Prize has scored Luna, 59.3% on ARC-AGI-2 against the 5.6 build's 59.5%, and has published nothing for Sol. OpenAI's own evaluations are all cost arguments: on AutomationBench 1.0.6 Sol at xhigh effort takes 33.2% at $0.27 a task, and the closest rival row is Claude Fable 5.1 with an Opus 5 fallback at 31.4%, so the real lead is 1.8 points rather than the 6.3 the Claude Opus 5 row implies. OpenAI says Sol makes about half as many mistakes as GPT-5.6 Sol, "approaching Astra-level reliability at much lower cost", and repeats that Astra is still its best model across the board. Nothing here moves a crown: Sol and Luna are cheaper, not better. There is no GPT-6 Terra, and OpenAI has not said whether one is planned. Full detail in our GPT-6 Sol and Luna breakdown.
Sep 22 Anthropic Claude Opus 5.5 takes the Intelligence Index at 57.6 and undercuts Opus 5 on price New
Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new Claude 5.5 family, and it is the largest single-day move this page has recorded. Artificial Analysis measures it at 57.6 on Intelligence Index v4.3.2, 4.3 points above Claude Fable 5.1 and the highest score it has published for any model, and it takes GDPval-AA v2.1 at 1846 against Fable 5.1's 1735, AA-Briefcase v1.1 at 1822 against 1678, Humanity's Last Exam at 61.4% against 59.1% and Artificial Analysis's own Terminal-Bench 4.0 run at 59.6% against GPT-6 Astra's 59.1% with it. Pricing goes down rather than up: $4 / $20 per 1M tokens against Opus 5's $5 / $25, cache reads 60% cheaper at $0.20 and cache writes at $5, with Anthropic's own testing putting a typical workload 40% cheaper than on Opus 5 and output more than 30% faster. A fast mode in Claude Code and the Claude Platform doubles the price for up to 2.5x the speed. On Anthropic's own harness the coding gap is wider still, 66.4% on Terminal-Bench 4.0 against the 57.9% OpenAI reports for Astra, plus 54.4% on FrontierCode v1.1 and 57.8% on CursorBench 4.0. Two things it does not win: Zapier's AutomationBench, where Astra scores 41.4% to its 40.0%, and Terminal-Bench-Science 0.1, where Astra scores 64.6% to its 58.7%. Read the margins with the caveat Anthropic prints itself, that the model "performs at the level of Claude Fable 5.1 on most work" and that "benchmark margins have become a less reliable guide to real-world differences". It is generally available on the Claude Platform as claude-opus-5-5 and on AWS, Google Cloud and Microsoft Azure, five-hour limits rise on Pro, Max and Team, and Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks. Because it matches Claude Mythos 5.1 in biology and cybersecurity it ships with Fable 5.1's safeguards and a new Life Sciences Verification Program for vetted labs. No Arena board has rated it yet. Arena has since ranked it first on WebDev and fourth on text. Full detail in our Claude Opus 5.5 breakdown.
Sep 21 SpaceXAI Grok 4.7 arrives on a larger base model at an unchanged $2 / $6, and costs 47% more per task New
SpaceXAI released Grok 4.7 on September 21, 2026 as its most capable model for coding and knowledge work. It uses a new and larger base model than Grok 4.6, with a longer reinforcement learning run weighted toward tasks that take many hours, and it was trained to understand the Grok Bot harness natively. The company publishes no parameter count. On its own figures it takes 46.3% on CursorBench 4.0 against Grok 4.6's 40.4% and GPT-5.6 Sol's 41.7%, 71.0% on DeepSWE v1.1 at high effort, 64.0% on EEBench, 38.0% on Terminal-Bench 4.0, 19.6% on the Harvey legal agent benchmark and 56.7% on HealthBench Professional. Independently, Artificial Analysis puts it at 46.3 on Intelligence Index v4.3.2 at high effort against Grok 4.6's 44.3, and at 1695 on GDPval-AA v2.1 and 1657 on AA-Briefcase v1.1, both above GPT-6 Astra's 1542 and 1569. The price does not move, $2 / $6 per 1M tokens with the whole request re-billed at $4 / $12 above 200,000 prompt tokens, and a fast variant runs at twice the output speed for twice the price. What does move is the cost per task: $2.73 against Grok 4.6's $1.86, because 4.7 emits roughly twice the output tokens getting to the same place, and xhigh effort buys no extra index points on either generation. On safety SpaceXAI reports the strongest refusal and jailbreak resistance it has tested, 62.4% on LatchBio's biosafety benchmark and 3.3% of risky dual-use prompts allowed through on its own HackerBench v0.3. Availability is developer-side: the Grok API, Cursor, Grok Build, third-party coding harnesses and cloud routers. The announcement says nothing about the consumer Grok app, which still serves Grok 4.6. Full detail in our Grok 4.7 breakdown.
Sep 21 Xiaomi Xiaomi's MiMo-V2.6-Pro is the highest-scoring open model ever measured, at $0.13 an index task under MIT New
Xiaomi published MiMo-V2.6-Pro on September 21, 2026 under plain MIT, with the weights on Hugging Face at XiaomiMiMo/MiMo-V2.6-Pro-RL on the day. It is a 1.02-trillion-parameter mixture of experts with 42 billion active and a 1M-token context. Artificial Analysis measures it, rather than estimating it, at 46.3 on Intelligence Index v4.3.2, which is the highest open score it has ever published: above GLM-5.3 at 44.8, above Kimi K3 at 43.6 and level with SpaceXAI's brand-new Grok 4.7. The cost is the other half of the story. It runs an Intelligence Index task for $0.13 at $0.435 / $0.87 per 1M tokens, which is cheaper per task than any other open model on this page, roughly half what GLM 5.3 Flash costs for four and a half points more. It also takes 34.8% on Terminal-Bench 4.0, 1673 on GDPval-AA v2.1, 49.4% on Humanity's Last Exam and 86.3% on AA-LCR, which puts it above Kimi K3 on three of those four. Two limits are worth stating plainly. No Arena board has rated it, so there is no human-preference evidence for it at all, and it was published four days ago, so it has no independent track record beyond Artificial Analysis. And 1.02 trillion parameters is a cluster, not a workstation, which is why GLM 5.3 Flash stays our recommendation for teams that actually want to host their own. It takes the open-weight crown on this page on measured score and cost. Arena has since rated it, 23rd on WebDev and 25th on text.
Sep 20 Alibaba Qwen-Image-2.1 is the best image model you can download, but its licence allows research use only New
Alibaba's Qwen team released Qwen-Image-2.1 on September 20, 2026, on Hugging Face, ModelScope and GitHub at once, and it is now the best image model you can download. It is one model for both generating and editing, with a 7B-parameter generator, down from about 20 billion in the original Qwen-Image. It makes transparent PNGs natively and edits with up to 10 reference images. On the independent boards it is the highest-ranked open-weight model on all four: 17th on Arena image editing at 1366.4, 18th on Arena text-to-image at 1227.6 and 18th on both Artificial Analysis boards. The catch is the licence. The 2.1 weights ship under the Qwen Research License for "research or evaluation purposes only", a change from the Apache 2.0 terms of earlier Qwen-Image releases, so it is not a free commercial option. For commercial work the older Apache 2.0 releases are the safe choice, though they rank far lower, Qwen-Image-2512 43rd on Arena text-to-image. Alibaba's hosted Qwen-Image-3.0-Pro, launched in July, ranks higher at 12th on Arena text-to-image, but it has no weights and costs $0.04 to $0.075 per image through the API. Full detail in our Qwen image generator guide.
Sep 15 TypeSafe TypeSafe leaves stealth with Jev, a model that returns typed decisions instead of text, at $0.042 per 1M tokens New
TypeSafe AI came out of two years of stealth on September 15, 2026 with $40 million in seed funding led by DCVC and a model that does not generate text at all. Jev is the first of what the company calls System One models: you send a block of state plus a fixed set of typed questions, and it answers all of them in one parallel pass, returning a choice, a score or a yes-no probability instead of a sentence. Because the answer space is defined before the call it cannot return a malformed result, which is where the launch post's line that Jev can't hallucinate comes from. It can still pick the wrong option, and TypeSafe's own documentation says calibration does not guarantee that an individual answer is correct. Pricing is $42 per billion input tokens, which is $0.042 per 1M, with output free, on a 64k context and text input only. The performance claims need care, because the company published four different ones in three days: from the press release's under 100 milliseconds and up to 100x faster to the homepage's 193.6x faster and 444.6x cheaper, each measured against a different rival, with its own footnote calling the big numbers the higher end of real world gains. Its workflow eval scores against the average of GPT-6 Astra and Fable 5.1 rather than ground truth, and it publishes no public benchmark scores at all by choice. Two independent early-access tests on September 18 measured roughly 6x cheaper and 8x faster against a cheap model with reasoning off, and the calibration held up. Nothing here moves a pick on this page: Jev has no consumer surface and is waitlisted early access. Full detail in our System One models explainer.
Sep 10 DeepSeek DeepSeek retires V4-Flash, replaces it with V4.1-Flash and cuts the Flash rate by about a third New
DeepSeek retired DeepSeek-V4-Flash on September 10, 2026 and put DeepSeek-V4.1-Flash in its place. The model name is now deepseek-flash; the legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names still resolve, but the models behind them have been retired and the requests are served by V4.1-Flash and billed at the Flash price. That price is a cut of roughly a third: $0.15 / $0.60 per 1M tokens off-peak and $0.30 / $1.20 at peak, against $0.22 / $0.66 and $0.44 / $1.32 for the build it replaces, with cache hits at $0.003 off-peak. The peak windows are unchanged, 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, with every other hour at half. V4-Pro is untouched at $0.66 / $1.98 off-peak, and DeepSeek has said it will keep serving it past its September 14 retirement date. The new model is a 552-billion-parameter mixture-of-experts on a causal encoder-decoder design that activates 8 billion parameters at prefill and 16 billion at decode, takes a 1M-token context, reads images natively, and went up on Hugging Face under MIT the same day. Artificial Analysis scores it 39.5 on Intelligence Index v4.3 against 34.5 for the V4-Flash 0731 build it replaces, at $0.27 an index task. It does not retake the price-performance pick: GLM 5.3 Flash scores 41.9 at $0.25 a task. Full detail in our DeepSeek pricing guide.
Sep 3 OpenAI GPT-6 Astra ships at $10 / $50, and within a day it is on Plus and Pro and leading the coding boards New
OpenAI launched GPT-6 Astra on September 3, 2026, settling the argument that ran all year: this is GPT-6, not another point release inside the GPT-5 line. Company president Greg Brockman told a press briefing "Welcome to the AGI era." The independent numbers are more measured than the framing. Artificial Analysis scores Astra 61 on Intelligence Index v4.1.1 at max effort, exactly level with GPT-5.6 Sol, five points behind Claude Fable 5.1 at 66 and two behind Claude Opus 5 at 63, and it charges $10 / $50 per 1M tokens against Sol's $4 / $20, which works out 75% more expensive per index task than the model it replaces. The stronger result is on the Coding Agent Index, where Astra in Codex scores 67, level with Claude Opus 5 and Claude Fable 5 in Claude Code and behind Claude Fable 5.1's 70, and it gets there on one third of Sol's tokens and one fifth of Opus 5's, which puts it on the cost-efficiency frontier. It also halves hallucination on AA-Omniscience, from 92% to 51% at max effort, while gaining four points of accuracy, and adds roughly 80 Elo on AA-Briefcase and six points on Humanity's Last Exam. Against that it drops about 80 Elo on GDPval-AA v2 and regresses two to three points on customer support, SciCode and long-context reasoning. Two caveats travel with the launch numbers: the 98.6% on ARC-AGI-3 is an action-efficiency score against a human baseline rather than a solve rate, measured on a harness OpenAI itself has shown can triple the result, and Epoch AI discloses that OpenAI funded FrontierMath and holds exclusive access to part of it. Availability is the real constraint. Astra is the first model OpenAI has ever designated Critical for cybersecurity under its Preparedness Framework, so approved defenders in the Daybreak program got it first. ChatGPT Business and Pro followed on September 4 and Plus a few hours later, with the API, AWS, Azure and the free tier still outstanding. Artificial Analysis has since retired the Coding Agent Index, and on its own Terminal-Bench 4.0 run Astra now leads outright, 60% to Claude Fable 5.1's 55%, which is what moved the coding crown on this page to Astra. Claude Opus 5.5 took that crown from it on September 22. Full detail in our GPT-6 Astra breakdown.
Sep 2 Alibaba Qwen3.8-Max-0902 takes Arena's WebDev board off Claude Opus 5 by three points, at $2 / $6 New
Alibaba shipped Qwen3.8-Max-0902 on September 2, 2026, and the naming is the first thing to get right: this is a post-training refresh of the Qwen3.8-Max that launched on August 3, not a new version, running the same 2.4 trillion parameters and the same 1M-token context window. Qwen says it was further post-trained on coding and Cowork-style work. The reason it matters is the board it moved. Arena announced the same day that Qwen3.8-Max-0902 debuted at #1 overall on Code Arena: WebDev with 1691 points, three points above Claude Opus 5 (Max), seventeen above Kimi K3 (Max) and twenty-two above the previous Qwen3.8-Max, taking #1 in the Data & Analytics and Consumer Product categories as well. That is the first time this year an Anthropic model has been displaced at the top of a vote-based coding board. Pricing is $2 / $6 per 1M tokens with cache hits at $0.17 explicit and $0.25 implicit, which Arena scores as a blended $5 per million and the best position on its Pareto frontier. Read it with three caveats. Three points is inside the noise on a vote-based board, Artificial Analysis had published no Intelligence Index for the 0902 snapshot at the time and has since scored it 45.4 on v4.3, and this is a hosted model on QwenCloud with no consumer chat front-end: the open weights Alibaba promised for Qwen3.8-Max in August still have not shipped, so nothing here changes our open-weight pick. Arena says Agent Arena scores are still to come. Qwen3.8-Max-0902 has since slipped to tenth on WebDev, and the coding category has moved to Claude Opus 5.5. Full detail in our Qwen 3.8 breakdown.
Sep 2 Meta Muse Spark 1.3, Intelligence Index 57 to 61 at an unchanged $1.25 / $4.25, winning coding and losing every agent row New
Meta released Muse Spark 1.3 on September 2, 2026, its fourth Muse Spark model in five months, in Muse Code and the Meta Model API. Artificial Analysis scores it 61 on Intelligence Index v4.1.1, up from 57 for 1.2 and 53 for 1.1, which is eight points in two months and the fastest climb anyone on this page is managing. Pricing did not move at all: $1.25 / $4.25 per 1M tokens on a 1M-token context window, with the Contributor tier still at $0.10 / $0.20 in exchange for letting Meta train on your prompts and completions. Two things about the launch need reading carefully. On Meta's own scorecard the model wins every coding row, DeepSWE v1.1 at 75.4 against Claude Opus 5's 74.0 and SWEAtlas CodeBase QnA at 59.4, ties GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8, and takes both MRCR long-context bands by a distance, 98.1 on the 512K to 1M band against 55.5 for 1.2. It also loses all six rows Meta groups under Agent, four to Opus 5 and two to GPT-5.6 Sol, and that half of the chart went almost unreported. The other catch is which model was measured: the benchmarked column is Muse Spark 1.3 (max), a limited preview for Meta's partners that scores 62, while the version anyone can call today is xhigh at 61, and Meta ran the 1.2 comparison column at xhigh rather than max. Meta shipped it to developers first, on Muse Code and the Meta Model API, and said a rollout to Facebook, Instagram and the Meta AI app would follow in the coming days, and the open-weights promise slipped again: August's pledge to release the Muse Spark 1.2 weights became an undated line about open weights coming soon. Full detail in our Muse Spark 1.3 breakdown.
Sep 2 Google Gemini 3.8 Flash, three points on the Intelligence Index at the same $0.75 / $3.75 New
Google made Gemini 3.8 Flash generally available on September 2, 2026, three weeks after 3.7 Flash and its third Flash release in 43 days. Every specification carries over unchanged: 1M input context, a 64K output limit, a March 2026 knowledge cutoff, the same modality list, and the same introductory $0.75 / $3.75 per 1M tokens that doubles to $1.50 / $7.50 on January 1, 2027. Only the scores moved, and they moved on every published row. DeepSWE v1.1 goes from 65.3% to 73.7%, Terminal-bench 2.1 from 85.8% to 89.4%, Terminal-bench 4.0 from 11.2% to 19.1% and BioMysteryBench Human Difficult from 43.5% to 56.5%. Artificial Analysis scores it 59 on Intelligence Index v4.1.1 against 56 for 3.7 Flash, and measures 304.6 output tokens per second against 279.4, with a slow 13.39-second time to first token that makes it a batch and agent model rather than an interactive one. Read the wins alongside the losses: Google's own comparison table gives 3.8 Flash 8 of 14 rows, and Claude Opus 5 still takes Terminal-bench 4.0 by 51.8% to 19.1%, OSWorld-2.0 by 75.4% to 59.0% and GDPVal-AA v2 by 1824 to 1545 on Elo. The DeepSWE figure republished across much of the launch coverage, 71.0%, is not the one in Google's PDF, which reads 73.7% against Opus 5's 74.0%. Developer access was live at announcement through the Gemini API, AI Studio, Antigravity and the Gemini Enterprise Agent Platform. Consumer access is reported for Google AI Pro and Ultra subscribers, but the Gemini Apps release notes carry no entry for it yet, and the free Gemini tier drops from 3.6 Flash to Flash-Lite on October 9, 2026. Full detail in our Gemini 3.8 Flash breakdown.
Sep 1 Anthropic Claude Fable 5.1 and Mythos 5.1, a new #1 on Artificial Analysis at 66 and a 75% cut to cache reads New
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, and they are one model in two safeguard configurations rather than two models. Fable 5.1 is generally available; Mythos 5.1 relaxes the biology and cybersecurity restrictions and is handed out by invitation, with no published eligibility criteria. Artificial Analysis scores Fable 5.1 at 66 on its v4.1.1 Intelligence Index at max effort, the highest score it has ever measured, ahead of Claude Opus 5 at 63, Claude Fable 5 at 62, and GPT-5.6 Sol and Grok 4.6 at 61. It also took the then-current Agentic Index at 61 against Opus 5's 59, the GDPval-AA v2 professional-deliverables board at 1853 against 1824, and Humanity's Last Exam at 59.1% against Fable 5's 55.5%. The effort setting matters more than usual here: the same model reads 58 at low, 60 at medium, 62 at high and 65 at xhigh, and those settings span an eleven-fold difference in output tokens, from 13.1 million to 143.7 million across the evaluation suite. Pricing is unchanged at $10 / $50 per 1M tokens, but cache reads drop 75% to $0.25 from Fable 5's $1.00, which is where the saving actually sits on long agentic runs. On Anthropic's own figures it reaches 52.6% on Terminal-Bench-Science 0.1 against Opus 5's 29.0%, and the 212-page system card raised the company's own alignment risk assessment from very low to low. Two things to weigh before switching: Anthropic still recommends starting with Opus 5 for most workloads at half the price, and at launch no Arena board had votes for either model. Fable 5.1 has since been rated: it is fourth on Arena WebDev, second on image-to-WebDev and sixth on the text board. Full detail in our Claude Fable 5.1 and Mythos 5.1 breakdown.
Pending Google Gemini 3.5 Pro, still unreleased, and Gemini 4 Argon has overtaken it Delayed
Gemini 3.5 Pro is still unreleased. Google announced it at I/O on May 19 alongside Gemini 3.5 Flash, but only Flash shipped, and Bloomberg reported on July 16, 2026 that it was months behind schedule. Google DeepMind's Koray Kavukcuoglu said on September 24 that the company "took a little bit of a step back" to focus on its Flash models and on Gemini 4, and on September 30 Google announced Gemini 4 Argon as its next flagship instead. Google has not said whether 3.5 Pro will ship at all. Use Gemini 3.8 Flash in the meantime if you are working through the API, or Gemini 3.6 Flash in the Gemini app, where from October 9, 2026 it needs Google AI Plus or above.
Category pick, October 2026

Best AI for Writing

1
Claude Fable 5.1
Best writing overall
→
2
Kimi K3
Creative fiction
→
3
Claude Sonnet 5.5
Free everyday

The best AI for writing is Claude Fable 5.1, and the case for it is consistency rather than a single first place. It is inside the top ten of all three prose boards: second on EQ-Bench Creative Writing v3 at 2162.0, third on LiveBench Language at 89.5 and tenth on Arena's creative-writing board. No other model places that well across all three. Gemini 4 Argon now leads Arena's creative-writing and text boards, at 1518.8 and 1525.2, but neither EQ-Bench nor LiveBench has rated it and only Google's Fairwind partners can use it. Claude Opus 5.5 is second on Arena's creative-writing board at 1516 but seventh on EQ-Bench and eighth on LiveBench Language; GPT-6 Astra leads EQ-Bench at 2173.3 but sits 39th on Arena's creative-writing board; Claude Fable 5 tops LiveBench Language at 90.7 and is third on Arena's creative-writing board, but eleventh on EQ-Bench. If your writing is a work deliverable rather than prose, the GDPval-AA v2.1 professional-deliverables board is a different question, and Claude Opus 5.5 leads it at 1866, with Claude Sonnet 5.5 27 points behind at 1839 and Claude Fable 5.1 at 1758. Claude Sonnet 5.5 is the free-tier value pick, since anyone can use it on claude.ai, though it scores 83.4 on LiveBench Language, sits only 42nd on Arena's creative-writing board and has no EQ-Bench rating yet. Claude Fable 5 is the pick if you weight Arena's human votes above the harness scores; it costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. Translation is a separate question with a separate winner, which we work through in our guide to the best AI for translation.

ModelBest ForStrengthWeaknessPrice (per 1M tokens)
Claude Fable 5.1Best writing overall#2 EQ-Bench Creative Writing (2162.0), #3 LiveBench Language (89.5) and #10 Arena creative writing, the best combined placing of any modelLeads none of the three outright; Gemini 4 Argon now leads Arena's creative-writing board, with Claude Opus 5.5 second$10 / $50
Claude Fable 5Strong on the human-vote boards#1 LiveBench Language (90.7), #3 Arena creative writing (1502.6), #3 Arena text (1504.3)#11 EQ-Bench; priciest option here$10 / $50
Kimi K3Creative fiction and voice#5 EQ-Bench Creative Writing (2082.3)Only #26 on Arena creative writing$3 / $15
Claude Sonnet 5.5Free everyday writingFree on claude.ai, 1M context; #2 GDPval-AA v2.1 (1839), 27 points behind Opus 5.583.4 LiveBench Language; #42 Arena creative writing; not yet rated on EQ-Bench$2 / $10
Claude Opus 5.5Arena's top creative writer you can use#2 Arena creative writing (1516) and #4 Arena text (1504); #1 GDPval-AA v2.1 (1866)#7 EQ-Bench and #8 LiveBench Language$4 / $20
Gemini 4 ArgonNot available yet#1 Arena creative writing (1518.8) and #1 Arena text (1525.2)Unrated on EQ-Bench and LiveBench; Fairwind partners only$2 / $10 intro, $4 / $20 standard
GPT-5.5Fact-anchored business writingDocumented factual-reliability gains over GPT-5.4Reasoning tiers now marked deprecated$5 / $30
Gemini 3.8 FlashBulk drafts at scaleIntelligence Index 40.9, 308.4 tok/s output, 1M contextAgent-tuned rather than prose-tuned; intro price doubles January 1, 2027$0.75 / $3.75
Runner-up and alternatives: Claude Fable 5 is the runner-up and the pick if you weight Arena's human votes, Claude Opus 5.5 is the highest-placed model you can use on Arena's creative-writing board and the leader for professional deliverables, Kimi K3 is the pick for distinctive fiction though it now sits fifth on EQ-Bench, Claude Sonnet 5.5 is the value pick and the one to use if you are not paying, and Gemini 3.8 Flash is the pick for bulk drafting. Gemini 4 Argon is the one to watch once Google opens access.
What changed this month

The writing crown stays with Claude Fable 5.1, but its margin on Arena narrowed. Google's Gemini 4 Argon, announced September 30, went straight to first on Arena's creative-writing board at 1518.8, pushing Claude Opus 5.5 to second and Claude Fable 5 to third, and Fable 5.1 slipped from seventh to tenth, so this card's claim is now the top ten of all three prose boards rather than the top seven. Argon cannot take the crown: EQ-Bench and LiveBench have not rated it, and only Google's Fairwind partners can use it. In September, Claude Opus 5.5 and Claude Sonnet 5.5 arrived, Sonnet 5.5 replaced Sonnet 5 as the free pick, and GPT-6.1 Sol entered LiveBench Language second, pushing Fable 5.1 to third there.

Category pick, October 2026

Best AI for Chat & Daily Assistant

1
GPT-5.6
ChatGPT's default
→
2
Claude Opus 5.5
#1 on the index
→
3
Claude Sonnet 5.5
Free Claude

The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #20 on Arena's text leaderboard at 1483.9, where Gemini 4 Argon now leads at 1525.2 and Claude Opus 5.5 is fourth at 1503.7. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (the 5.6 tiers: Luna $0.20 / $1.20, Terra $2 / $12, Sol $4 / $20 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.

ModelBest ForStrengthWeaknessPrice
GPT-5.6Everyday chat, ChatGPT's defaultThe assistant most people can open; Terra matches GPT-5.5 at ~half cost#20 on Arena text; scheming flagged by METRFree / $20/mo Plus; API $0.20 / $1.20 to $4 / $20
Claude Opus 5.5Highest-scoring conversation#1 on the Intelligence Index (57.6) and #4 on Arena text (1503.7)Not on the Free plan; usage-limited on Pro$20/mo Pro, $4 / $20 API
Gemini 4 ArgonNot available yet#1 on Arena text (1525.2); 15% hallucination rate, the lowest among leading modelsFairwind partners only; not in the Gemini app$2 / $10 API intro, once released
Claude Sonnet 5.5Free Claude chatFree on claude.ai; #2 on the Intelligence Index (56.0), behind only Claude Opus 5.5Only #45 on Arena text; runs at Medium effort by default in the appsFree / $20/mo Pro, $2 / $10 API
GPT-5.5 InstantHallucination-sensitive daily work52.5% fewer hallucinated claims vs 5.3 InstantReasoning tiers now marked deprecated$20/mo Plus; API $5 / $30
Gemini 3.6 FlashFast, cheap, multimodalThe free Gemini tier’s model until October 9 (then AI Plus), 1M context, #22 on Arena textWeaker on hardest reasoning; 3.8 Flash outscores it at the same priceFree / $0.75 / $3.75 API
Fello AIAll the top models, one appChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPadRouted via app, not direct$9.99/mo
Runner-up and alternatives: Claude Opus 5.5 is the runner-up and the top scorer on Artificial Analysis's index, Claude Sonnet 5.5 is the pick for free Claude chat, Gemini 3.6 Flash is the runner-up for fast and cheap (free until October 9, then Google AI Plus), and Grok 4.6 is the niche pick for live-news days. Fello AI is the natural pick if you want the top models in one Mac and iOS app for $9.99/month instead of juggling subscriptions.
What changed this month

GPT-5.6 keeps the chat pick on reach, and nothing about who can open what changed: GPT-5.6 Luna is still ChatGPT's free default, and the GPT-6 models still run in ChatGPT Work and Codex rather than in Chat. The preference board moved again. Gemini 4 Argon, announced September 30, took first on Arena's text board at 1525.2, and Claude Opus 5.5, which led it a week ago, is now fourth at 1503.7, about a point behind Claude Opus 4.6 and Claude Fable 5. Opus 5.5 keeps the second podium place as the top scorer on Artificial Analysis's index; Argon stays off the podium because only Google's Fairwind partners can use it and it is not in the Gemini app. In September, Claude Sonnet 5.5 became the free Claude model and took the third podium place from Claude Opus 5.

Category pick, October 2026

Best AI for Images

1
ChatGPT Images 2.5
Text in images
→
2
MAI-Image-2.6
#3 AA image editing
→
3
Grok Imagine Image 2.0
Fewest restrictions

The best AI for image generation is ChatGPT Images 2.5, which OpenAI made the default on every ChatGPT tier on September 8 and which leads all four image boards we track. Its two models finish first and second on each: GPT-Image-2.5 Sunburst, built for precision, leads Arena text-to-image at 1424.6 and Arena image editing at 1523.6, and Artificial Analysis's AA-Image-T2I v2.0 at 1197 and AA-Image-Editing v2.0 at 1182, with the faster GPT-Image-2.5 Flare second on all four. Read the Sunburst-to-Flare gap on Arena text-to-image as provisional, because each rests on about 17,000 votes against GPT Image 2's 93,000. MAI-Image-2.6 is the runner-up, third on Artificial Analysis's editing board and fourth on Arena's text-to-image board at 1333.3. Third is SpaceXAI's Grok Imagine Image 2.0, released in August and available in Grok Imagine and the xAI API: it is fourth on Artificial Analysis's text-to-image board and fourth on Arena image editing, ahead of Meta's Muse Image on three of the four boards. Google's Nano Banana 2.1, released on October 6, is now the highest-ranked Google model: fifth on Arena text-to-image at 1328.1, in the same rank band as MAI-Image-2.6 and Grok Imagine Image 2.0, and sixth on Arena editing, though Artificial Analysis has not rated it yet. The best image model you can download is Alibaba's Qwen-Image-2.1, released on September 20 and now the highest-ranked open-weight model on all four boards, 18th on Arena image editing and on both Artificial Analysis boards and 19th on Arena text-to-image, but its licence allows research and evaluation only, so it is not a free commercial option. Our full breakdown of what shipped is in ChatGPT Images 2.5.

ModelBest ForStrengthWeaknessPrice
ChatGPT Images 2.5Images with readable text#1 and #2 on all four boards: Arena text-to-image (1424.6 / 1398.5), Arena editing (1523.6 / 1481.0), Artificial Analysis text-to-image (1197 / 1191) and editing (1182 / 1162)Its Arena text-to-image lead rests on about 17,000 votes per modelDefault on every ChatGPT tier
MAI-Image-2.6Editing an image you already have#3 on Artificial Analysis image editing (1137) and #4 on Arena text-to-image (1333.3)Public preview, not in Copilot or Bing yet; lost the Artificial Analysis editing lead to ChatGPT Images 2.5MAI Playground / Microsoft Foundry
Grok Imagine Image 2.0Fewest restrictions, Spicy Mode#4 on Artificial Analysis text-to-image (1155) and #4 on Arena image editing (1439.1); sixth on Arena text-to-imageOnly #10 on Artificial Analysis image editing$30/mo SuperGrok; xAI API
Nano Banana 2.1Google's everyday image model#5 Arena text-to-image (1328.1) and #6 Arena image editing (1428.1), the highest of any Google model; half Nano Banana 2's API price per imageNot yet rated by Artificial Analysis; input tokens cost three times as much as on Nano Banana 2Gemini app / AI Studio; $0.0336 per 1K image
Seedream 5.0 ProMultilingual text + region editing10+ languages incl. Arabic RTL, lasso and layer editingAround tenth on the independent boards; copyright cloudBytePlus / Magnific
Midjourney v8Stylized art, illustrationAesthetic baseline most artists preferWeaker on text in image$10-$120/mo
Qwen-Image-2.1Best image model you can downloadHighest-ranked open-weight model on all four boards: #18 Arena image editing (1368.7), #19 Arena text-to-image (1223.3) and #18 on both Artificial Analysis boards; 7B generator, transparent PNGs, up to 10 reference imagesQwen Research License, research or evaluation only, no commercial useFree download (Hugging Face)
Qwen-Image-3.0-ProDense, text-heavy layouts#13 Arena text-to-image (1255.0) and #14 Artificial Analysis text-to-image (1090); prompts up to 4.5k tokens, text in 12 languagesHosted only, no weights; not rated on Arena image editingQwen Chat; $0.04-$0.075 per image in the API
Muse ImageImage editing, Meta ecosystem#6 on Artificial Analysis image editing (1118) and #7 on its text-to-image boardOff the podium; #9 on Arena text-to-image and #8 on Arena editingMeta AI app
Runner-up and alternatives: MAI-Image-2.6 is the runner-up overall and the pick for editing an image you already have, and Grok Imagine Image 2.0 is third and still the only frontier model that allows Spicy Mode adult content. Muse Image is the pick inside Meta's apps, Nano Banana 2.1, which replaced Nano Banana 2 on October 6, is the pick inside Google's apps, and Reve 2.1 remains the pick for layout, typography and native 4K, seventh on Arena text-to-image. If you want to run a model on your own hardware, Qwen-Image-2.1 scores best of any open model, but only for research and evaluation; for commercial work the older Apache 2.0 Qwen-Image releases are the safe choice, though they rank far lower, Qwen-Image-2512 44th on Arena text-to-image.
What changed this month

ChatGPT Images 2.5 keeps the image crown and still holds first and second on all four boards we track. Since the last update Google released Nano Banana 2.1 on October 6, and Arena has already ranked it fifth on text-to-image and sixth on editing, the highest of any Google model, inside the same rank band as MAI-Image-2.6 and Grok Imagine Image 2.0. It takes Nano Banana 2's table row rather than a podium place: it trails MAI-Image-2.6 on both Arena boards, splits the two with Grok Imagine Image 2.0, and Artificial Analysis has not rated it yet. Its arrival moves Grok Imagine Image 2.0 to sixth and Reve 2.1 to seventh on Arena text-to-image. Earlier this month Grok Imagine Image 2.0 took third place from Muse Image, and Artificial Analysis moved both its image boards to v2.0. In September, Images 2.5 replaced GPT Image 2 at the top on September 8.

Category pick, October 2026

Best AI for Video

1
MiniMax H3
#1 Arena image-to-video
→
2
Gemini Omni 1.1 Flash
#1 Arena text-to-video
→
3
Wan 3.0
#1 AA text-to-video

The best AI for video generation is MiniMax H3, which MiniMax released on July 31 with open weights, and the case for it is that it places near the top everywhere rather than winning one board outright. It leads Arena's image-to-video board at 1495.4, ahead of Gemini Omni 1.1 Flash at 1488.5 and Alibaba's Wan 3.0 at 1479.8; on Artificial Analysis's image-to-video board it is second at 1181, behind only MiniMax H3 Max, a model Fal built on top of it; and it is fourth on Artificial Analysis's text-to-video board at 1137. Its weak spot is Arena's text-to-video board, where it is eighth at 1460.3. It makes 2K clips of 4 to 15 seconds with native stereo audio, from $0.13 a second at 2K, and the weights are on Hugging Face under MiniMax's own community licence, without the 2K upscaler and with an application required in the US, EU, UK and South Korea. Gemini Omni 1.1 Flash, our pick until this update, still leads Arena's text-to-video board at 1516.2, in a statistical tie with Gemini Omni Flash at 1512.9, and it is the pick for long takes: scene extension in 10-second steps to 40 seconds, first and last frame control and 4K output, from about $0.03 a second at 360p to $0.30 at 4K. But Artificial Analysis has now rated it, and it comes eighth on text-to-video at 1113 with no image-to-video rating at all. Wan 3.0 leads Artificial Analysis's text-to-video board at 1156 and is third on Arena's image-to-video board, but only sixth on Arena's text-to-video board. If you need longer takes inside Google's tooling, Veo 3.1 still runs in the Gemini app, AI Studio and Vertex AI with native audio and 1080p output.

ModelBest ForStrengthWeaknessPrice
MiniMax H3Best AI video overall#1 Arena image-to-video (1495.4), #2 Artificial Analysis image-to-video (1181), #4 Artificial Analysis text-to-video (1137); 2K clips of 4-15s with native stereo audio#8 on Arena text-to-video (1460); the open weights exclude the 2K upscaler and need an application in the US, EU, UK and South Korea$0.13/sec at 2K
Gemini Omni 1.1 FlashLong takes and 4K#1 Arena text-to-video (1516), #2 Arena image-to-video (1488.5); 40-second scene extension, 4K outputOnly #8 on Artificial Analysis text-to-video (1113); no Artificial Analysis image-to-video rating$0.03-$0.30/sec; Gemini API / AI Studio / Flow
Gemini Omni FlashCheaper 10-second clips#2 Arena text-to-video (1513), #3 Artificial Analysis image-to-video (1178); conversational editingCaps at 10-second generations~$0.10/sec; Gemini app / AI Studio
Wan 3.0Video from documents and slides#1 Artificial Analysis text-to-video (1156), #3 Arena image-to-video (1479.8); 2-30s, up to 1080p, Omni-Reference inputAPI-only, no published weights; #6 on Arena text-to-video$0.05-$0.20/sec
Dreamina Seedance 2.5ByteDance challenger#3 Artificial Analysis text-to-video (1143), #4 Arena image-to-video (1477.4)#7 Arena text-to-video (1474); ByteDance ecosystem, limited Western accessDreamina / BytePlus
Muse VideoMeta ecosystem video#9 on Arena text-to-video at 1456Newest of the group, thin toolingMeta AI app
Veo 3.1Longer production clipsNative audio, 1080p, strong physics consistency#12 on Arena video, not the quality leaderGoogle AI Pro / Ultra
Kling 3.0 / 3.0 TurboFast iteration at lower costNative 4K, 60fps, 15-second clips; Turbo shipped June 17Outside the top 16 on Arena text-to-videoFrom $10/mo
Luma Ray 3Photoreal scenesStrong realism for landscapesSmaller communityFree / from $9.99/mo
Runner-up and alternatives: Gemini Omni 1.1 Flash is the runner-up and the pick for text-to-video and long takes, with 40-second scene extension and 4K output, and Gemini Omni Flash is the cheaper Google option for 10-second clips. Alibaba's Wan 3.0 is third and the one to use when the source material is a document, a spreadsheet or a slide deck. Dreamina Seedance 2.5 is the ByteDance pick, and Veo 3.1 is the pick inside the Gemini app and Vertex AI when you want native audio at 1080p. OpenAI retired the Sora 2 consumer app on April 26, 2026 and the developer API was scheduled to close on September 24, 2026. We track the retirement dates for older models across every provider in a separate running list.
What changed this month

The video crown moves from Gemini Omni 1.1 Flash to MiniMax H3. Since the last update Artificial Analysis has rated Omni 1.1 Flash for the first time, and on its text-to-video board, now at version 2.0, it comes eighth at 1113, behind Wan 3.0 at 1156 and MiniMax H3 at 1137. With H3 still first on Arena's image-to-video board and second on Artificial Analysis's, behind only a model built on it, H3 now places better across the four boards we track. Omni 1.1 Flash keeps first place on Arena's text-to-video board, still a statistical tie with Gemini Omni Flash, so it stays on the podium as the text-to-video pick, and Wan 3.0 takes third place from Dreamina Seedance 2.0. Earlier this month FLUX 3 Video and Grok Imagine Video 1.5 entered Arena's text-to-video board third and fourth. In September, Gemini Omni 1.1 Flash held first place on Arena's text-to-video board and MiniMax H3 took first place on its image-to-video board.

Category pick, October 2026

Best AI for Coding

1
Claude Opus 5.5
#1 Arena WebDev
→
2
Claude Sonnet 5.5
Value pick, #1 Terminal-Bench
→
3
GPT-6 Astra
#1 image-to-WebDev

The best AI for coding is Claude Opus 5.5, which Anthropic shipped on September 22, though it no longer leads every board, and the model that took two of them is its own cheaper sibling. Claude Sonnet 5.5, released September 28, leads Artificial Analysis's own Terminal-Bench 4.0 run at 63.6% against Opus 5.5's 59.6% and LiveBench Coding at 91.4 against 89.3. Opus 5.5 keeps the crown on the rest of the evidence. It leads Arena's WebDev board at 1815, ahead of GPT-6 Astra at 1788 and Claude Sonnet 5.5 at 1786. It scores 71.7 on LiveBench Agentic Coding, second only to DeepSeek V4.1-Flash, where Sonnet 5.5 manages 56.3, and on Anthropic's own harnesses it beats Sonnet 5.5 on FrontierCode v1.1, 54.4% to 52.1%, and on CursorBench 4.0, 57.8% to 55.5%. Price does not settle it the way the token rates suggest. Sonnet 5.5 is $2 / $10 per 1M tokens against Opus 5.5's $4 / $20, but at max effort it spends so many tokens that an Intelligence Index task costs $7.67 against Opus 5.5's $5.98. Its value is at lower effort: at xhigh it scores 51.9 on the index for $2.75 a task and at high 46.7 for $1.12. GPT-6 Astra keeps Arena's image-to-WebDev board at 1733 and matches Opus 5.5's 59.6% on Terminal-Bench 4.0 at xhigh, at $10 / $50, and Claude Fable 5.1 is fifth on WebDev and second on image-to-WebDev. GPT-6.1 Sol is the cheap OpenAI option: fourth on WebDev and 56.1% on Terminal-Bench 4.0 for $0.72 an index task, though LiveBench puts it well outside the top ten on Coding, at 80.7. Gemini 4 Argon leads Google's own DeepSWE v1.1 table at 77.9%, but Artificial Analysis measures 57.1% on Terminal-Bench 4.0, Arena puts it ninth on WebDev, and it is not available outside Google's Fairwind program. Among open weights, Xiaomi's MiMo-V2.6-Pro takes 34.8% on Terminal-Bench 4.0 under plain MIT at $0.13 an index task, GLM-5.3 is still the strongest open model on that harness at 41.9%, and Kimi K3 is the highest-placed open model on Arena WebDev, thirteenth at 1658, on only 12.6% of the harness. Artificial Analysis has retired both its Coding Index and its Coding Agent Index, so the two composite coding rankings this page used to quote no longer exist.

ModelBest ForStrengthWeaknessPrice (per 1M tokens)
Claude Opus 5.5Best coding overall#1 Arena WebDev (1815) and #1 Intelligence Index (57.6, v4.3.2); 71.7 LiveBench Agentic Coding, 15 points above Sonnet 5.5Second to Sonnet 5.5 on Terminal-Bench 4.0 (59.6% to 63.6%) and LiveBench Coding (89.3 to 91.4)$4 / $20
Claude Sonnet 5.5Value pick for coding#1 Terminal-Bench 4.0 (63.6%) and #1 LiveBench Coding (91.4), #3 Arena WebDev (1786); 51.9 on the index for $2.75 a task at xhigh56.3 LiveBench Agentic Coding; at max effort $7.67 an index task, more than Opus 5.5$2 / $10
GPT-6 AstraTop of Arena image-to-WebDev#1 Arena image-to-WebDev (1733), #2 WebDev (1788); 59.6% Terminal-Bench 4.0 at xhigh2.5x Opus 5.5's price; behind Sonnet 5.5 on Terminal-Bench and well behind both Claude models on LiveBench Coding$10 / $50
Claude Fable 5.1Highest composite score before Opus 5.5#5 Arena WebDev and #2 image-to-WebDev; Intelligence Index 53.4; AA-Briefcase 167652.0% Terminal-Bench 4.0; 2.5x the price of Opus 5.5$10 / $50
Claude Opus 5Superseded by Opus 5.549.0% Terminal-Bench 4.0 and #7 Arena WebDev and #3 image-to-WebDevCosts more per token than Opus 5.5 and scores seven points lower$5 / $25
Claude Fable 5Hardest long-horizon agentic workIntelligence Index 49.6; 42.4% Terminal-Bench 4.0; 1M contextThe priciest per index task on this table at $8.75; #6 on Arena image-to-WebDev$10 / $50
Qwen3.8-Max-0902Brief holder of Arena's WebDev board#11 Code Arena: WebDev (1670); Intelligence Index 45.4; 2.4T parameters, 1M contextQwenCloud API only with no consumer front-end; lost the WebDev lead within days$2 / $6
Kimi K3Highest-placed open model on Arena#13 Arena WebDev (1658), the best open placement on that board; 2.8T/104B activeOnly 12.6% on Terminal-Bench 4.0; 1.56 TB to self-host$3 / $15
GPT-5.6 SolOpenAI's cheaper flagshipIntelligence Index 47.0; 39.9% Terminal-Bench 4.0Superseded by GPT-6 Sol, which scores higher at half the price; eval-gaming flagged by METR$4 / $20
GPT-6 SolCheap GPT-6 coding in CodexIntelligence Index 47.6 (v4.3.2) and 43.9% Terminal-Bench 4.0, both above GPT-5.6 Sol, at half its price#8 Arena WebDev (1689); in Codex and ChatGPT Work only, not in Chat$2 / $10
GPT-6.1 SolCheapest top-tier model per taskIntelligence Index 51.8 (v4.3.2), #4 Arena WebDev (1758) and 56.1% Terminal-Bench 4.0 at $0.72 an index task, with cached input at $0.1080.7 on LiveBench Coding, well outside the top ten; in Codex and ChatGPT Work only, not in Chat$2 / $10
Gemini 4 ArgonNot available yet77.9% DeepSWE v1.1 on Google's own figures; Intelligence Index 52.6 at $1.99 an index task57.1% Terminal-Bench 4.0 and #9 Arena WebDev (1680); Fairwind partners only$2 / $10 intro, $4 / $20 standard
Muse Spark 1.3Cheap agentic codingIntelligence Index 48.1 at $1.25 / $4.25, and $1.60 per index task33.3% on Terminal-Bench 4.0; the coding wins come from Meta's own chart, measured on a partners-only max variant$1.25 / $4.25
Grok 4.7Cheap value coder46.3% CursorBench 4.0 and 71.0% DeepSWE v1.1 on SpaceXAI's own figures; Intelligence Index 46.324.7% on Terminal-Bench 4.0; prompts of 200K tokens and up re-bill the whole request at double$2 / $6
Gemini 3.8 FlashAgent coding at scaleIntelligence Index 40.9; 308.4 output tokens per second19.7% on Terminal-Bench 4.0; intro price doubles January 1, 2027$0.75 / $3.75
GLM 5.3 FlashCheapest serious coder you can host32.8% Terminal-Bench 4.0 at $0.25 an index task; MIT, 320B/18B activeGLM-5.3 reaches 41.9% on the same harness; no consumer productOpen weights (MIT)
MiMo-V2.6-ProHighest open Intelligence Index46.3 on v4.3.2 and 34.8% Terminal-Bench 4.0 at $0.13 an index task; MIT, 1.02T/42B activeOnly #24 on Arena WebDev (1618), eleven places behind Kimi K3$0.435 / $0.87
Runner-up and alternatives: Claude Sonnet 5.5 is the runner-up and the value pick, leading Terminal-Bench 4.0 and LiveBench Coding at half Opus 5.5's token price as long as you keep it below max effort, GPT-6 Astra is third and still tops Arena's image-to-WebDev board, GPT-6.1 Sol is fourth on WebDev at a fifth of Astra's token price, and Claude Opus 5 is now superseded by a model that is both cheaper and better. Among open weights, Xiaomi's MiMo-V2.6-Pro is the leader on score and on cost, GLM-5.3 is the strongest on Artificial Analysis's Terminal-Bench 4.0, Kimi K3 is the highest-placed open model on Arena's WebDev board, and GLM 5.3 Flash is the pick for teams hosting their own, at $0.25 an index task under plain MIT. Inside IDEs, Cursor with Claude is still the most popular pairing and Claude Code is the natural pick if you live in the terminal.
What changed this month

Claude Opus 5.5 keeps the coding crown. Since the last update Claude Sonnet 5.5 has entered Arena's WebDev board third at 1786, two points behind GPT-6 Astra, which pushes GPT-6.1 Sol to fourth and Claude Fable 5.1 to fifth, and Gemini 4 Argon, which entered at the start of the month, is now ninth. Argon leads Google's own DeepSWE v1.1 table at 77.9%, but on Artificial Analysis's Terminal-Bench 4.0 it scores 57.1%, behind Sonnet 5.5, Opus 5.5 and GPT-6 Astra, and nobody outside Google's Fairwind program can use it. In September the crown moved twice, from Claude Fable 5.1 to GPT-6 Astra and then to Opus 5.5, and Claude Sonnet 5.5 took Terminal-Bench 4.0 and LiveBench Coding as the value pick.

Category pick, October 2026

Best AI for Creativity

1
Grok 4.7
Fewest guardrails
→
2
Claude Fable 5
Highest-quality prose
→
3
Kimi K3
Fiction and voice

The best AI for unfiltered, on-trend creative work is Grok 4.7, and we want to be exact about why. This pick is about the product line, not the prose quality. Grok carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. Read the availability carefully: SpaceXAI shipped Grok 4.7 on September 21 to the Grok API, Cursor, Grok Build, third-party harnesses and cloud routers, and its announcement says nothing about the consumer app, which still serves Grok 4.6 to SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, though the picture is more mixed than it was. EQ-Bench has now rated Grok 4.7 and puts it eighth at 2006.7, inside its top ten and far above Grok 4.5 in 52nd, but Arena's human voters put 4.7 only 64th on creative writing, behind Grok 4.6 in 52nd and the older Grok 4.20-beta1 in 24th. If you are picking on output quality alone, Gemini 4 Argon now leads Arena's creative-writing board, though almost nobody can use it yet, with Claude Opus 5.5 second and Claude Fable 5 third, while EQ-Bench puts GPT-6 Astra first at 2173.3, Claude Fable 5.1 second and Kimi K3 fifth. Choose Grok for what it will let you make, not for how well it writes.

ModelBest ForStrengthWeaknessPrice
Grok 4.7Unfiltered, opinionated, on-trendFewest content restrictions, native real-time X grounding#64 on Arena creative writing, though #8 on EQ-Bench (2006.7)$2 / $6 API; $30/mo SuperGrok app, still on 4.6
Claude Fable 5Highest-quality creative prose#3 Arena creative writing, #1 LiveBench LanguageCautious guardrails on edgy briefs$10 / $50 API
Kimi K3Fiction and distinctive voice#5 EQ-Bench Creative Writing at 2082.3Only #26 on Arena creative writing$3 / $15
Claude Opus 5Long-form structured creativityHolds long threads and self-edits; Intelligence Index 50.8, now behind six newer modelsMost cautious of the group$20/mo Pro; $5 / $25 API
Gemini 3.1 ProMultimodal creativeStrong text, image and video chainQuotas inside the Gemini appFree / $2.00-$4.00 API in
Grok Imagine (Spicy Mode)NSFW / adult creativeMost permissive image generationNiche use case$30/mo SuperGrok
Runner-up and alternatives: Claude Fable 5 is the runner-up and the right pick if quality matters more than freedom, Kimi K3 is the pick for fiction, and Claude Opus 5 is the pick for creative projects that run across many turns. For adult creative work, Grok Imagine Spicy Mode is still the only frontier-grade option.
What changed this month

Grok keeps this pick on Grok 4.7, and nothing about the product's permissiveness or its live X access changed, which is what the pick rests on. Arena's creative-writing board moved around it: Gemini 4 Argon, announced September 30, went in at first, Claude Opus 5.5 is now second and Claude Fable 5 third, and Grok 4.7 edged up from 72nd to 64th. EQ-Bench has not changed and still puts 4.7 eighth. The split surface still matters for this category: 4.7 is on the API, Cursor and Grok Build, while the $30/month Grok app that the permissiveness argument is really about still serves 4.6. In September, Grok 4.7 replaced Grok 4.6 as SpaceXAI's flagship on September 21 and both creative-writing boards rated it.

Category pick, October 2026

Best AI for Accuracy & Research

1
Claude Fable 5.1
Highest factual accuracy
→
2
GPT-6.1 Sol
Best value, $0.72/task
→
3
Claude Opus 5.5
Best when wrong costs most

The best AI for accuracy and research is Claude Fable 5.1, which holds the highest factual accuracy Artificial Analysis has measured, 67% on AA-Omniscience. That is the exact column this category is decided on. Claude Opus 5.5 beat Fable 5.1 on almost everything else when it arrived on September 22, including Humanity's Last Exam at 61.4% to 59.1% and the hallucination-penalised AA-Omniscience index at 46.4 to 43.5, but on raw factual accuracy it reads 66% to Fable 5.1's 67%. Treat that one-point gap as a tie and read the two together: Opus 5.5 is the better bet when a wrong answer is worse than no answer, Fable 5.1 when you want the most facts right. The value pick is now GPT-6.1 Sol, which answers 62% correctly with an AA-Omniscience index of 41.5 at $2 / $10 per 1M tokens and $0.72 an index task, and scores 98.5% on ARC Prize's ARC-AGI-1 for $0.06 a task. That beats Gemini 3.1 Pro, our value pick until this update, on every accuracy measure Artificial Analysis publishes: it reads 55% raw accuracy and an index of 31.9. Gemini 3.1 Pro keeps its case where the answer has to be current, because it grounds answers in Google Search and is in the Gemini app (Google AI Pro from October 9), while GPT-6.1 Sol runs in ChatGPT Work, Codex and the API rather than in ChatGPT's chat window. Gemini 4 Argon is the model to watch: Artificial Analysis measures a 15% hallucination rate, the lowest among leading models, because it declines rather than guesses, but it answers only 50% correctly and only Google's Fairwind partners can use it. On grounded search specifically, Arena's search leaderboard is led by OpenAI: GPT-5.6 Sol tops it at 1257, with Claude Fable 5 fifth and Gemini 3.1 Pro grounding ninth. On novel reasoning the crown sits elsewhere, with GPT-6 Astra leading ARC-AGI-2 and ARC-AGI-3. Claude Opus 5 holds no accuracy crown because Artificial Analysis measures its hallucination rate at 61%.

ModelBest ForKey BenchmarkWeaknessPrice
Claude Fable 5.1Highest measured factual accuracy67% factual accuracy on AA-Omniscience, the highest Artificial Analysis has measured, one point above Claude Opus 5.5No cheap tier; unrated on Arena's search leaderboard; Opus 5.5 leads Humanity's Last Exam and the AA-Omniscience index$10 / $50
Claude Opus 5.5When a wrong answer is worse than none#1 AA-Omniscience index (46.4) and #1 Humanity's Last Exam (61.4%); 66% raw accuracyOne point behind Fable 5.1 on raw accuracy$4 / $20
GPT-6.1 SolBest value62% raw accuracy and an AA-Omniscience index of 41.5 at $0.72 an index task; 98.5% ARC-AGI-1 at $0.06 a task54% hallucination rate; in ChatGPT Work and Codex, not in Chat$2 / $10
Gemini 4 ArgonFewest hallucinations, once available15% hallucination rate on AA-Omniscience, the lowest among leading modelsOnly 50% raw accuracy; Fairwind partners only$2 / $10 intro
Gemini 3.1 ProGrounded search; Gemini app on AI Pro from Oct 9Native Google Search grounding; 98% ARC-AGI-1 at $0.52/task55% raw accuracy on AA-Omniscience, below GPT-6.1 Sol; #9 on Arena search$2.00-$4.00 / $12.00-$18.00 (tiered)
Claude Fable 5Grounded search#5 on Arena's search leaderboard, behind GPT-5.6 SolNo single cheap tier$10 / $50 (Fable 5)
GPT-5.6 SolNovel reasoning#4 ARC-AGI-2 at 92.5%, behind GPT-6 Astra (95%), GPT-6.1 Sol (94.2%) and Claude Opus 5.5 (93.3%)Scheming flagged by METR$4 / $20
Claude Opus 5Hardest unseen problems30.2% ARC-AGI-3 on the standard harness, second to GPT-6 Astra (ARC Prize)Artificial Analysis measures a 61% hallucination rate$5 / $25
Qwen 3.7 MaxFrontier accuracy at value pricing92.4 GPQA Diamond, 200 free requests/dayAPI-only, no chat front-end$1.25 / $3.75 promo; $2.50 / $7.50 list
Claude Opus 4.6Honesty under pressure#1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5Superseded as a flagshipLegacy Anthropic model
Runner-up and alternatives: Claude Opus 5.5 is the runner-up and the pick when a wrong answer costs more than no answer, GPT-6.1 Sol is the value pick at $0.72 an index task, Gemini 3.1 Pro is the pick when answers need Google Search grounding or you want it inside the Gemini app (Google AI Pro from October 9), Claude Opus 4.6 sweeps the honesty-under-pressure board, and Qwen 3.7 Max is the value pick at the frontier.
What changed this month

The accuracy crown stays with Claude Fable 5.1, at 67% raw factual accuracy against Claude Opus 5.5's 66%, and the value pick moved. GPT-6.1 Sol now beats Gemini 3.1 Pro on every accuracy measure Artificial Analysis publishes, 62% raw accuracy against 55% and an AA-Omniscience index of 41.5 against 31.9, and ARC Prize scores it 98.5% on ARC-AGI-1 for $0.06 a task against Gemini 3.1 Pro's 98% for $0.52, so it takes the value card. Gemini 3.1 Pro stays in the table for grounded search and the Gemini app, where it needs Google AI Pro from October 9. Gemini 4 Argon, announced September 30, brings the lowest hallucination rate Artificial Analysis has measured among leading models, 15%, but it answers only half its questions correctly and is not yet available, so it changes nothing at the top. Two older lines are corrected as well: Claude Opus 5's hallucination rate now reads 61% rather than 50%, and it no longer leads ARC-AGI-3. In September, Opus 5.5 took the hallucination-penalised index and Humanity's Last Exam, which is why this block now splits the two Claude models rather than crowning one outright.

Category pick, October 2026

Best AI for Problem Solving

1
GPT-6 Astra
Reasoning leader
→
2
Claude Opus 5
Agentic chains
→
3
GPT-6.1 Sol
Value pick, $2 / $10

The best AI for hard problem solving is GPT-6 Astra, which took the reasoning boards off GPT-5.6 Sol within days of launching. It leads LiveBench Reasoning at 92.65 and ARC-AGI-2 at 95%, the highest anyone has scored against that board's 100% human panel, and sits fourth on LiveBench Mathematics at 96.81, behind Claude Opus 5.5 at 97.08, Claude Fable 5.1 at 97.01 and GPT-6.1 Sol at 96.83. It also reports 97.6% on FrontierMath Tier 4 v2, a number worth carrying with Epoch AI's disclosure that OpenAI funded the benchmark and holds exclusive access to part of it. The catch is price and reach: $10 / $50 per 1M tokens against GPT-6.1 Sol's $2 / $10, and it needs a paid ChatGPT plan. GPT-6.1 Sol, released September 29, is the value alternative and very nearly a co-leader: 92.63 on LiveBench Reasoning, 0.02 behind Astra, and 94.2% on ARC-AGI-2, 0.8 points behind, for $0.25 a task against Astra's $1.12. Astra keeps the crown because it still leads both boards and ARC-AGI-3, where it scores 62.7% on the standard harness against GPT-6.1 Sol's 52.7%. Gemini 4 Argon leads Arena's math category, but on only a few hundred votes, and neither LiveBench nor ARC Prize has rated it. GPT-5.6 Sol, the old value pick, is now fifth on LiveBench Reasoning and fourth on ARC-AGI-2 at 92.5%. Qwen 3.7 Max is the budget pick for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, with 200 free model requests a day. Claude Opus 5 remains the alternative for long agentic reasoning chains, though Astra has taken ARC-AGI-3 from it as well.

ModelBest ForKey BenchmarkWeaknessPrice
GPT-6 AstraHardest math, science and reasoning#1 LiveBench Reasoning (92.65), #1 ARC-AGI-2 (95%), #1 ARC-AGI-3Needs a paid ChatGPT plan; 5x GPT-6.1 Sol's price$10 / $50
Claude Opus 5Long agentic reasoning chains1662 on AA-Briefcase v1.1 and 1724 on GDPval-AA v2.1, fourth among distinct models on both; 30.2% ARC-AGI-3Astra now leads ARC-AGI-3; 61% hallucination rate; now a legacy model$5 / $25
GPT-5.6 SolValue STEM flagship#4 ARC-AGI-2 (92.5%); #5 LiveBench Reasoning (91.65), Mathematics 96.20FrontierMath still unpublished; scheming flagged by METR$100/mo ChatGPT Pro; API $4 / $20
GPT-6.1 SolValue pick, near-Astra reasoning#2 LiveBench Reasoning (92.63), #2 ARC-AGI-2 (94.2%), #3 LiveBench Mathematics (96.83); Intelligence Index 51.8 at $0.72 an index task52.7% on ARC-AGI-3 against Astra's 62.7%; in ChatGPT Work and Codex, not in Chat$2 / $10
Qwen 3.7 MaxCompetition math on a budget97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/dayAPI-only$1.25 / $3.75 promo; $2.50 / $7.50 list
Claude Fable 5Math inside a coding workflow#3 on Arena's math subcategory (1522), behind Gemini 4 Argon and Claude Opus 5; 96.0 LiveBench MathematicsPriciest option here$10 / $50
GLM 5.3 FlashOpen-weight problem solvingIntelligence Index 41.8 under MIT, almost eight points above GLM-5.2 at less than half the size; 320B/18B active, 1M contextNeeds a multi-GPU server, not a workstation; GLM 5.3 scores higher on Z.ai's own figures and is downloadable since August 28, but at more than twice the size and under a bespoke licenceOpen weights (MIT)
Runner-up and alternatives: GPT-6.1 Sol is the runner-up and the value pick at $2 / $10, Claude Opus 5 is the natural pick for long-chain agentic reasoning, Qwen 3.7 Max is the budget pick, and GLM 5.3 Flash is the open-weight pick. Our dedicated guide to the best AI for math covers the task-by-task split and the free options.
What changed this month

GPT-6 Astra keeps the problem-solving crown, by less than a point. ARC Prize has now scored GPT-6.1 Sol, and it takes 94.2% on ARC-AGI-2, second to Astra's 95% and ahead of Claude Opus 5.5's 93.3%, for under a quarter of Astra's cost per task. Together with its 92.63 on LiveBench Reasoning, 0.02 behind Astra, that makes GPT-6.1 Sol a near co-leader, but Astra still leads both boards and ARC-AGI-3, so the crown does not move. GPT-5.6 Sol drops to fourth on ARC-AGI-2. Gemini 4 Argon, announced September 30, leads Arena's math category on a few hundred votes but has no LiveBench or ARC Prize score yet, and Claude Fable 5, which led that Arena category, is now third. In September, Astra took this crown from GPT-5.6 Sol after its September 3 launch, and LiveBench Mathematics passed to Claude Opus 5.5.

Category pick, October 2026

Best AI Agent

1
Gemini Spark
Cloud-resident
→
2
Claude Cowork
Desktop-resident
→
3
OpenAI dots
Always-on, runs on GPT-6 Astra

The best AI agent is Gemini Spark if you want the agent in the cloud and Claude Cowork if you want it on your desktop. These are joint picks rather than a first and a second: Spark runs in a Google Cloud VM, so it keeps working while your laptop is shut, wired into Gmail, Docs and, since early September, Google Photos; Cowork runs on your Mac or Windows machine and drives your local apps and your screen. Spark is no longer the only agent that keeps working in the cloud. SpaceXAI's Grok Bot (August 11) gives each account a persistent cloud computer that signs into your apps, from a $20/month Cursor Pro plan; Meta's Muse (September 8) runs errands from a cloud VM of its own and is free up to 100 million tokens a week; and OpenAI dots (September 29) run on GPT-6 Astra with their own cloud computer and browser, from the $100/month Pro plan, which does not include them in the EEA, UK or Switzerland. No independent board rates these products against each other, so the choice is about where the work has to happen rather than which one scores higher. Price is no longer a tiebreaker either: Spark now reaches the $19.99/month Google AI Pro tier in the US and more than 160 other countries, and Cowork is included at no extra charge on every paid Claude plan from $20/month. ChatGPT Codex Mobile (May 14) is the alternative for coding-agent work. Read the full Gemini Spark vs Claude Cowork comparison.

AgentBest ForWhere It RunsStrengthPrice
Gemini Spark24/7 cloud tasks, Workspace workflowsGoogle Cloud VM (always-on)Always-on, deep Workspace integrationFrom $19.99/mo Google AI Pro
Claude CoworkDesktop, app-driving, design + codeYour Mac/Windows desktopDrives local apps, sees your screen$20/mo Claude Pro
ChatGPT Codex MobileCoding agent on phoneOpenAI cloud + iOS/AndroidApprove diffs and redirect work from phoneIncluded in ChatGPT plans
Grok BotBrowser jobs across logged-in appsOne persistent SpaceXAI cloud computer shared by all your BotsSigns into apps with your own logins and keeps going with your laptop closedFrom $20/mo Cursor Pro, or a linked SuperGrok plan
OpenAI dotsOngoing responsibilities, not one-off tasksOpenAI cloud computer and browserRuns on GPT-6 Astra; keeps working with your computer off; reachable in ChatGPT, Slack and by voiceFrom $100/mo ChatGPT Pro (not EEA/UK/CH); Business Premium
Meta MuseFree personal errandsA Muse Secure VM in Meta's cloudFree up to 100M tokens a week; a separate Sentinel agent approves what it sends outFree; $20/mo Power, $100/mo Maximum
Runner-up and alternatives: Gemini Spark and Claude Cowork are joint picks split by where the agent runs, not a first and a second. OpenAI dots are the pick if you already pay for ChatGPT Pro and want an agent on OpenAI's flagship, Grok Bot is the pick for browser jobs across your logged-in apps, from a $20/month Cursor Pro plan, Meta's Muse is the one to try for free, and ChatGPT Codex Mobile is the pick for coding agents. Google AI Ultra at $99.99 or $199.99 a month now buys Spark higher usage limits rather than access, which starts on the $19.99 Pro tier.
What changed this month

Gemini Spark (cloud) and Claude Cowork (desktop) remain joint picks, because this category crowns products and no new agent product shipped since the last update. The model layer moved. Claude Sonnet 5.5 has entered Arena's Agent Arena third, behind Claude Fable 5.1, which still tops it with 17.6% confirmed success, and Claude Opus 5.5, and it has taken the lead on Artificial Analysis's AA-Briefcase v1.1 at 1823, with Opus 5.5 second at 1807. Opus 5.5 still leads GDPval-AA v2.1 at 1866, with Sonnet 5.5 second at 1839. Gemini 4 Argon, announced September 30, leads Artificial Analysis's AutomationBench at 77.5% and has a 15.4% confirmed-success rate on the Agent Arena, above Opus 5.5's 14.1%, but sits tenth there overall and is not yet available outside Google's Fairwind program. Among open models on the Agent Arena, Kimi K3 is now the highest at fourteenth, ahead of DeepSeek V4.1-Flash at seventeenth and Tencent's Hy4 Preview at nineteenth. In September, Grok Bot, Meta's Muse and OpenAI dots joined Gemini Spark as always-on cloud agents, and Spark reached the $19.99/month Google AI Pro tier.

Fello AI running on Mac, iPhone and iPad
All the top AI models, one app

The leading models like ChatGPT, Claude and Gemini, together on Mac, iPhone and iPad.

Free to start, 4.7★ across 27,000+ reviews.

Download Fello AI

Use case guide, October 2026

Best AI for Students

1
GPT-5.6 Luna
Essays & research
→
2
Gemini 3.6 Flash
STEM + PDFs
→
3
Claude Sonnet 5.5
Writing & editing

The best AI for students is GPT-5.6 Luna inside ChatGPT for general coursework and Gemini 3.6 Flash inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Sonnet 5.5 as the free alternative for essay editing. Most students don't need to pay, and the free tier improved on August 6: GPT-5.6 Luna became the default for ChatGPT Free and Go, with unlimited text chats and a Think button for harder questions, though file uploads and image tools stay capped. On September 22 OpenAI added GPT-6 Luna for Free and Go users in the ChatGPT desktop app, a current-generation model on the free tier from launch day, which Artificial Analysis now scores 38.1 on its index against 37.3 for GPT-5.6 Luna. Gemini 3.6 Flash is what the free Gemini app serves until October 9, when free accounts drop to Flash-Lite, Claude Sonnet 5.5, released September 28, is free on claude.ai, and DeepSeek V4 is free on DeepSeek's chat site. Luna is the cheapest member of the GPT-5.6 family rather than the strongest, so reach for the Think button or switch to GPT-5.5 when an essay needs more care. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's paid flagship and GPT-6 Astra now reports 97.6% on FrontierMath Tier 4 v2 against GPT-5.5 Pro's verified 39.6%, though Astra is not on ChatGPT's free tier; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).

TaskBest ModelWhyFree?Alternative
Essays & courseworkGPT-5.6 LunaFree default in ChatGPT since August 6; unlimited text chats, plus GPT-6 Luna in the desktop app since September 22YesClaude Sonnet 5.5 (free Claude)
STEM problem-solvingGPT-5.6 Sol / Qwen 3.7 MaxPaid STEM flagship; Astra reports 97.6% FrontierMath Tier 4 v2 / 97.1 HMMT 2026 FebPro paid / Qwen API paidGemini 3.6 Flash (free until Oct 9, then AI Plus)
Research & accuracyGemini 3.1 Pro98% ARC-AGI-1 at $0.52/task, native Google Search groundingUntil Oct 9 (Gemini app), then AI ProClaude Opus 5
Writing editingClaude Sonnet 5.5Free on claude.ai since September 28 and second only to Opus 5.5 on the Intelligence Index; Claude Fable 5.1 is the quality leaderYes (Claude free)Claude Fable 5.1
Multimodal study (PDFs, slides, images)Gemini 3.6 Flash1M context via AI Studio; Gemini app needs AI Plus from Oct 9Free in AI Studio; app until Oct 9NotebookLM (Google)
Runner-up and alternatives: Claude Sonnet 5.5 (free) is the runner-up for essay writing and editing. Gemini 3.6 Flash (free until October 9, then Google AI Plus) is the runner-up for multimodal study and PDF ingestion. DeepSeek V4 is the runner-up for problem-solving on a strict zero-cost budget.
What changed this month

One change for students is scheduled: from October 9 the free Gemini app drops from Gemini 3.6 Flash to Flash-Lite, so 3.6 Flash needs Google AI Plus. Otherwise GPT-5.6 Luna is still the free ChatGPT default and Claude Sonnet 5.5 is still free on claude.ai. Gemini 4 Argon, announced September 30, is not available to students at all yet, and Google has said nothing about bringing it to the free Gemini app or the Google AI Pro plan. In September, Claude Sonnet 5.5 replaced Sonnet 5 as the free writing and editing pick, and OpenAI added GPT-6 Luna for Free and Go users in the ChatGPT desktop app, a cheaper model that now scores slightly higher than GPT-5.6 Luna.

Use case guide, October 2026

Best AI for Work & Professionals

1
GPT-5.6
Daily knowledge work
→
2
Claude Opus 5.5
Coding & writing
→
3
Gemini 3.1 Pro
Research & briefings

The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5.5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark is the always-on cloud agent that lives inside Google Workspace, and since July 30 it reaches the $19.99/month Google AI Pro tier as well as Google AI Ultra; OpenAI dots, Grok Bot and Meta's Muse are the alternatives, covered in the agents section above.

Use CaseBest ModelKey StatPriceAlternative
Daily knowledge workGPT-5.6ChatGPT's default since July 9; the assistant most people can open$20/mo ChatGPT PlusClaude Opus 5.5
ChatGPT Work & CodexGPT-6.1 SolIn ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu since September 29, at $2 / $10 in the API$20/mo ChatGPT PlusGPT-6 Luna for volume work
Coding (proprietary)Claude Opus 5.5#1 Arena WebDev (1815) and #1 Intelligence Index (57.6), at $4 / $20$20/mo Claude ProClaude Sonnet 5.5, which leads Terminal-Bench 4.0 at half the token price
Coding (cost-effective)Claude Sonnet 5.5#1 Terminal-Bench 4.0 (63.6%) and #1 LiveBench Coding (91.4); $1.12 an index task at high effort$2 / $10Qwen3.8-Max-0902; DeepSeek V4.1-Flash
Research & briefingsGemini 3.1 Pro98% ARC-AGI-1 at $0.52/task, Google groundingGoogle AI Pro / UltraClaude Opus 5.5
Hard math, physics, finance modellingGPT-6 Astra#1 LiveBench Reasoning and ARC-AGI-2 (95%); reports 97.6% FrontierMath Tier 4 v2$100/mo ChatGPT ProGPT-5.6 Sol; Qwen 3.7 Max
Always-on agent workflowsGemini SparkAlways-on cloud agent inside Google WorkspaceFrom $19.99/mo Google AI ProClaude Cowork
Live news, X-context creativeGrok 4.7Intelligence Index 46.3 + native X grounding; the Grok app still serves 4.6$30/mo SuperGrokGemini 3.1 Pro
All-in-one consolidationFello AIChatGPT + Claude + Gemini + Grok + DeepSeek$9.99/moPay each vendor separately
Runner-up and alternatives: for most professional teams, Claude Opus 5.5 is the runner-up to GPT-5.6 for daily work and the coding leader at $4 / $20, with Claude Sonnet 5.5 the cheaper coding pick that leads Terminal-Bench 4.0 and GPT-6 Astra the pick on Arena's image-to-WebDev board. Gemini 3.1 Pro is the runner-up for research-heavy roles, and Gemini Spark is the pick if you can put a cloud agent to work on long tasks inside Google Workspace, with OpenAI dots the alternative for ChatGPT Pro subscribers.
What changed this month

No row in this guide changes hands. Gemini 4 Argon, announced September 30, leads the Vals Index of finance, legal and tax work and Artificial Analysis's AutomationBench, which makes it the model to watch for professional work, but only Google's Fairwind partners can use it, with paid API customers and Google AI Ultra subscribers next and no date. In September, Claude Opus 5.5 took the proprietary coding row at a lower price than Opus 5, Claude Sonnet 5.5 took the cost-effective coding row, and GPT-6.1 Sol took the ChatGPT Work and Codex row on September 29.

Pricing Comparison

AI Model Pricing in October 2026

From $0 free tiers to $199.99/month Google AI Ultra. The price that matters most on this table is Claude Opus 5.5 at $4 / $20. Since September 22 the highest-scoring model any independent evaluator has measured costs less than the Claude Opus 5 it replaced ($5 / $25), and cache reads fall 60% to $0.20. Claude Fable 5.1 and GPT-6 Astra both list at $10 / $50 and both score below it.

The busiest rate is now $2 / $10, shared by Claude Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol. GPT-6.1 Sol is the cheapest route to a top-six Intelligence Index score, at $0.72 an index task. With Sonnet 5.5 the effort setting decides the bill: an index task costs $1.12 at high effort but $7.67 at max, more than Opus 5.5's $5.98. Google's Gemini 4 Argon also starts at an introductory $2 / $10 ($4 / $20 standard), but only Google's Fairwind partners can buy it so far.

Token price and task cost can drift apart. Grok 4.7 lists at $2 / $6, the same as Grok 4.6, yet an index task costs $2.73 against $1.86, because 4.7 writes roughly twice the output. Mistral Large 4 (public preview since October 6) lists at $1.36 / $4.18, though Mistral's own model page shows half that, and it is verbose enough that an index task costs $1.13 for a score of 38.4. At the frontier the value pick is Meta's Muse Spark 1.3 at $1.25 / $4.25, or $1.60 an index task for 48.1.

At the cheap end, GPT-6 Luna is the cheapest model per task on this page: $0.10 / $0.50, or $0.07 an index task. Among open models sold by the token, GLM 5.3 Flash stays the pick at $0.15 / $0.50 ($0.25 an index task for 41.8), ahead of DeepSeek V4.1-Flash at $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak ($0.27 for 39.5). Xiaomi's MiMo-V2.6-Pro is cheaper still at $0.13 an index task for 46.3, but only if you host its 1.02 trillion parameters yourself. For a deeper breakdown see our full AI Pricing Comparison Guide.

ModelInput (per 1M)Output (per 1M)ContextFree Access?
GPT-6 Astra$10.00$50.001,050,000 (922,000 max input)In Chat on Pro, Business and Enterprise since September 4; Plus gets it in ChatGPT Work and Codex rather than Chat; no free tier; live in the API
GPT-6 Sol$2.00$10.001,050,000 (922,000 max input)ChatGPT Work & Codex on Plus, Pro, Business, Enterprise and Edu since September 22; API; not in Chat
GPT-6.1 Sol$2.00$10.001,050,000 (922,000 max input)ChatGPT Work & Codex on Plus, Pro, Business, Enterprise and Edu since September 29; API; not in Chat; cache reads $0.10
GPT-6 Luna$0.10$0.501,050,000 (922,000 max input)Free and Go in the ChatGPT desktop app; ChatGPT Work & Codex on paid plans; API; not in Chat
GPT-5.5$5.00$30.001M (400K in Codex)ChatGPT Free; API paid
GPT-5.5 Pro$30.00$180.001MChatGPT Pro from $100/mo ($200 and $500 higher-usage tiers)
GPT-5.6 Sol$4.00$20.00Not publishedLive in ChatGPT, Codex & API (July 9); promotional rate through at least November 21, 2026
GPT-5.6 Terra$2.00$12.00Not publishedLive in ChatGPT, Codex & API (July 9)
GPT-5.6 Luna$0.20$1.20Not publishedLive in ChatGPT, Codex & API (July 9)
Claude Opus 5.5$4.00$20.001MClaude Pro/Max/Team; API paid; cache reads $0.20
Claude Opus 5$5.00$25.001MLegacy model at Anthropic; Pro/Max/API
Claude Opus 4.8$5.00$25.001MLegacy model at Anthropic; Pro/Max/API
Claude Fable 5.1$10.00$50.001MCache reads $0.25; in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits. Mythos 5.1 bills identically, by invitation only
Claude Fable 5$10.00$50.001MPermanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits
Claude Sonnet 5.5$2.00$10.001MFree on claude.ai; Pro/Max/Team; API paid; cache reads $0.20
Claude Sonnet 5$2.00$10.001MLegacy model since September 28; the fallback for flagged cyber requests; API paid
Claude Sonnet 4.6$3.00$15.001MAPI paid (superseded by Sonnet 5)
Gemini 3.1 Pro$2.00 (≤200K) / $4.00 (>200K)$12.00 (≤200K) / $18.00 (>200K)1MLimited Gemini app; API paid
Gemini 4 Argon$2.00 intro / $4.00 standard ($0.10 cached, intro)$10.00 intro / $20.00 standard1M (1M output)Fairwind Program partners only; paid API and Google AI Ultra next, no date
Gemini 3.8 Flash$0.75 intro / $1.50 from Jan 1, 2027$3.75 intro / $7.50 from Jan 1, 20271MAI Studio, Antigravity, Gemini Enterprise + paid API; in the Gemini app reported for AI Pro/Ultra
Gemini 3.7 Flash$0.75 intro / $1.50 from Jan 1, 2027$3.75 intro / $7.50 from Jan 1, 20271MAI Studio, Android Studio, Antigravity + paid API; in the Gemini app via Spark only (AI Pro/Ultra, excludes EEA/UK/CH/Nigeria)
Gemini 3.6 Flash$0.75 intro / $1.50 from Jan 1, 2027$3.75 intro / $7.50 from Jan 1, 20271MFree Gemini app default until Oct 9, then AI Plus; AI Studio; free API tier + paid API
Gemini 3.5 Flash-Lite$0.30$2.501MAI Studio; free API tier + paid API
Qwen3.8-Max-0902$2.00 ($0.17 explicit cache hit, $0.25 implicit)$6.001M (128K output)QwenCloud API only; no consumer chat front-end, no open weights
Qwen 3.7 Max$1.25 promo / $2.50 list$3.75 promo / $7.50 list1M200 free requests/day; API paid beyond that
MiniMax M3$0.30 (50% off $0.60)$1.20 (≤512K)1MOpen weights; hosting costs apply
LongCat-2.0Provider-dependentProvider-dependent1MOpen weights (MIT); hosting costs apply
NVIDIA Nemotron 3 UltraProvider-dependentProvider-dependent1MOpen weights (OpenMDW); hosting costs apply
Qwen 3.5 (open-weight)Self-host / TogetherSelf-host / Together1MOpen weights; hosting costs apply
Nex-N2-ProSelf-host / providersSelf-host / providers1MOpen weights (Apache 2.0); hosting costs apply
Rio 3.5 Open 397BSelf-host / providersSelf-host / providers1MOpen weights (MIT); hosting costs apply
Grok 4.3$1.25$2.501MFree consumer plan; API paid
Grok 4.6$2.00 (<200K) / $4.00 (≥200K)$6.00 (<200K) / $12.00 (≥200K)500KSpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel, Cloudflare
Grok 4.7$2.00 (<200K) / $4.00 (≥200K)$6.00 (<200K) / $12.00 (≥200K)500KSpaceXAI API, Cursor, Grok Build, harnesses and cloud routers; not in the Grok app
Muse Spark 1.3$1.25 ($0.10 Contributor tier)$4.25 ($0.20 Contributor tier)1MPaid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount
Kimi K3$3.00 ($0.30 cache-hit)$15.001MFree basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License)
Mistral Large 4$1.36 ($0.68 on Mistral's model page)$4.18 ($2.09)512K (1M per Mistral)Public preview API since October 6; open weights promised by the end of October, no licence named yet
Gemini Omni Flash (video)$1.50$17.50 (video output)10-second clipsGemini app / Flow; AI Studio + API
DeepSeek V4-Pro$0.66 off-peak / $1.32 peak ($0.022 cache-hit off-peak)$1.98 off-peak / $3.96 peak1MDeepSeek Chat free; API paid, peak/off-peak tiers since August 16
DeepSeek V4.1-Flash$0.15 off-peak / $0.30 peak ($0.003 cache-hit off-peak)$0.60 off-peak / $1.20 peak1MDeepSeek Chat free; API paid, peak/off-peak tiers; replaced V4-Flash on September 10
Kimi K2.7 CodeProvider-dependentProvider-dependent256KOpen weights; hosting costs apply
GLM-5.2Provider-dependentProvider-dependent1MOpen weights; hosting costs apply
GLM 5.3$1.40 ($0.26 cache-hit)$4.401MZ.ai API, GLM Coding Plan and ZCode; weights on Hugging Face since August 28 under the bespoke GLM 5.3 License
GLM 5.3 Flash$0.15 ($0.03 cache-hit)$0.501MZ.ai API, GLM Coding Plan, OpenRouter; MIT weights on Hugging Face; the launch promotion ended September 9
Tencent Hy4 Preview$0.834 ($0.042 cache-hit)$2.5011M+Open weights (Apache 2.0); Tencent Cloud TokenHub, OpenRouter, WorkBuddy, CodeBuddy
Qwen3.8-Flash-NextProvider-dependentProvider-dependent262K (to 1M)Open weights (Qwen Community License 1.0); hosting costs apply
MiMo-V2.6-Pro$0.435$0.871MOpen weights (MIT); hosting costs apply
ERNIE 5.1China-region pricingChina-region pricing256KBaidu free tier
Gemini Spark (agent)Not API-pricedNot API-priced1M (Gemini base)Google AI Pro $19.99, AI Ultra $99.99 or $199.99/mo
Fello AI (aggregator)Routed via appRouted via appModel-dependent$9.99/mo, free tier available

The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. The GPT-6 tiers carry the same surcharge past 272K input tokens, applied to the whole request, and their cached input reads are 90% off: $1.00 for Astra, $0.20 for GPT-6 Sol and $0.01 for Luna, with GPT-6.1 Sol at 95% off, $0.10.

Grok 4.6 works the same way but bites harder: at 200,000 prompt tokens and above, SpaceXAI re-bills the entire request at the higher rate rather than only the overage, so crossing the line by a thousand tokens doubles the cost of the whole call. DeepSeek is the other rate card that needs reading twice: since August 16 both V4 models bill at peak rates from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays, and at exactly half that in every other hour, so the same job can cost twice as much depending on when you run it.

If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.

Open-Weight Models

Best Open-Weight Models in October 2026

The best open-weight model in October 2026 is Xiaomi's MiMo-V2.6-Pro, published on September 21 under plain MIT: 1.02 trillion parameters with 42 billion active and a 1M-token context. Artificial Analysis scores it 46.3 on Intelligence Index v4.3.2, the highest open score it has ever published, ahead of GLM-5.3 at 44.8 and Kimi K3 at 43.6, at just $0.13 an index task. The catch: Arena's human voters rank it only 24th on WebDev and 26th on text.

Two other open models still lead somewhere. Kimi K3 is the highest-placed open model on Arena, 13th on WebDev and 14th on the Agent Arena. GLM-5.3 is the strongest open model on Terminal-Bench 4.0 at 41.9%, though on a bespoke Z.ai licence rather than MIT.

For teams running their own weights, the practical pick is still GLM 5.3 Flash (Z.ai, MIT): 41.8 on v4.3.2 from 320B total and 18B active. It is a separate model on a newly trained base, not a trimmed GLM 5.3, which is why it ships under MIT. A 1.02-trillion-parameter model is not a workstation proposition, and the K3 download alone is about 1.56 TB. For the cheapest tier, DeepSeek-V4.1-Flash replaced V4-Flash on September 10: MIT, natively multimodal, 39.5 on v4.3.2 and $0.15 / $0.60 off-peak.

Three more are worth watching. Mistral Large 4 (public preview since October 6) is the highest-scoring model from the US or Europe that promises open weights, at 38.4, but you cannot download it yet: Mistral says the weights follow by the end of October and has not named a licence. Alibaba's Qwen3.8-Flash-Next, 125B with only 6B active, previews the Qwen4 architecture and scores 39.8. Tencent's Hy4 Preview, 770B total and 49B active under Apache 2.0, has no Artificial Analysis rating but sits 17th on Arena WebDev.

ModelBest ForKey BenchmarkContext / LicenseWhere To Run
MiMo-V2.6-ProHighest open Intelligence Index ever measuredII 46.3 (v4.3.2), above GLM-5.3 and Kimi K3, at $0.13 an index task; 34.8% Terminal-Bench 4.0; GDPval-AA v2.1 1686; 1.02T/42B active1M / MITHugging Face (XiaomiMiMo), providers, self-host
Kimi K3Highest-placed open model on ArenaII 43.6 (v4.3.2); #13 Arena WebDev, the best open placement there; #14 Agent Arena; 2.8T/104B active1M / Kimi K3 LicenseHugging Face (96 shards, 1.56 TB), Moonshot API, providers
GLM 5.3 FlashBest open model most teams can actually runII 41.8 (v4.3.2), on the intelligence-versus-cost Pareto frontier at about $0.25 an index task; 320B/18B active, natively multimodal1M / MITHugging Face, Z.ai API ($0.15/$0.50), OpenRouter
GLM-5.2Fallback with an independent track recordII 33.7 (v4.3.2); #26 Arena WebDev (1,605.0), #20 Agent Arena; 744B/40B active1M / MITZ.ai, Hugging Face, OpenRouter
GLM 5.3Open since August 28, but on a bespoke licenceII 44.8 (v4.3.2), the second-highest open score we found; same 744B base as GLM-5.2, post-trained only, published checkpoint counted at 753B; CyberGym 84.5% and Terminal-Bench 3.0 28.3 (vendor figures)1M / GLM 5.3 License (MaaS above $10B revenue needs a Z.ai security review)Hugging Face, Z.ai API ($1.40/$4.40), GLM Coding Plan, ZCode
DeepSeek V4.1-FlashBest open value if you host it yourselfII 39.5 (v4.3.2) against 34.3 for the V4-Flash 0731 build it replaces; 552B, 8B active at prefill and 16B at decode1M / MITDeepSeek API ($0.15/$0.60 off-peak, $0.30/$1.20 peak since September 10), local
Tencent Hy4 PreviewLargest permissively licensed model yet770B/49B active; 2.99/4.00 on Tencent's own 203-task expert panel vs Kimi K3 2.94 and GLM 5.3 2.92 (vendor); no Artificial Analysis rating; #17 Arena WebDev (1633), #19 Agent Arena1M+ / Apache 2.0Hugging Face (BF16 and FP8), Tencent Cloud TokenHub, OpenRouter
Qwen3.8-Flash-NextPreview of the Qwen4 architecture125B total plus a 51B N-gram embedding, 6B active; 91.7 GPQA Diamond and 62.5 SWE-Bench Pro (vendor); II 39.8 (v4.3.2); #16 Arena WebDev (1637)262K to 1M / Qwen Community License 1.0Hugging Face, providers, self-host
LongCat-2.0Frontier open coder trained on Chinese chipsII 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active1M / MITHugging Face, GitHub, OpenRouter
MiniMax M3Cheap frontier-class multimodalII 44 (v4.1), 59% SWE-Bench Pro, multimodal1M / license TBDHugging Face, API $0.30/1M (50% off)
Nex-N2-ProStrongest open coding scoreII 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B activeQwen-based / Apache 2.0Hugging Face, providers, self-host
Kimi K2.7 CodeStrongest commercially-licensed open coder+21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active256K / Modified MITHugging Face, DeepInfra, providers
DeepSeek V4-ProAgentic real-world workII 44 (v4.1), 1.6T/49B active1M / MITDeepSeek API ($0.66/$1.98 off-peak, $1.32/$3.96 peak since August 16), local
Hy3Newest permissive-licence entrantII 41 (v4.1); #53 on Arena WebDev (1,507.4)Apache 2.0Hugging Face, providers, self-host
InklingThinking Machines' first open modelII 41 (v4.1), agentic 32.3, released July 15Open weightsHugging Face, providers, self-host
Inkling SmallSame family at a quarter the sizeII 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio inApache 2.0Hugging Face (BF16 and NVFP4), providers, self-host
NVIDIA Nemotron 3 UltraNVIDIA-tuned, fully permissive licenseII 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active1M / OpenMDWOpenRouter, Hugging Face, AWS (8x B200 self-host)
Qwen 3.5 (397B / 17B active)Multimodal, fast decode88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v61M / openTogether, OpenRouter, local
Qwen3.6-35B-A3BEfficient open agentic coder (3B active)86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active262K (to 1M YaRN) / Apache 2.0Hugging Face, OpenRouter, local
Qwen3.6-27BLaptop-runnable dense coder87.8 GPQA Diamond, dense 27B, multimodal256K / Apache 2.0Local Mac/PC, Hugging Face, OpenRouter
Rio 3.5 Open 397BQwen 3.5 fine-tune, multilingual reasoning70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5397B/17B active / MITHugging Face, providers, self-host
Llama 4 MaverickMeta-line flagship17B active / 400B total paramsLlama 4 licenseMeta cloud, Hugging Face, local
NVIDIA Nemotron 3 Nano OmniEdge / low-powerMultimodal, very small footprintCompact / openLocal, NVIDIA tool

Licensing matters as much as raw score here, and August widened the gap between the two groups. Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM 5.3 Flash (MIT), MiMo-V2.6-Pro (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Hy4 Preview (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use with no revenue test.

The rest carry conditions worth reading before you build on them: MiniMax M3 and MiniMax H3 ship under their own community licences, H3 additionally requiring an application from the US, EU, UK and South Korea; Kimi K3 sits on a custom licence with a revenue threshold for anyone reselling it as a service; GLM 5.3 requires a Z.ai security review of any model-as-a-service operator above $10 billion in revenue; and the Qwen Community License 1.0 on Qwen3.8-Flash-Next requires a separate licence from Qwen to run a model-as-a-service or an AI work-assistant business, though internal use is exempt. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.

How We Evaluate

Benchmarks, Prices, and Hands-On Use

Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 4.0 board, covering the Intelligence Index, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.

Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We rank on measured scores: a category crown moves as soon as a model out-scores the holder on the boards that define that category, without waiting for the vote-based boards to catch up, and a model that is not yet widely available is flagged on its card rather than denied one. We re-check ranks as well as figures, because a score can stay correct while the board around it moves. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.

Fello AI running on Mac, iPhone and iPad
Not sure which model to pick?

Fello AI brings the top AI models together in one lightweight native app.

Switch between them anytime, no separate subscriptions.

Download Fello AI
Frequently Asked Questions

Common Questions

What is the best AI model right now in October 2026?
It depends on the task, but Anthropic's two September releases still sit at the top of most boards. On overall benchmark score, Claude Opus 5.5 (September 22) is #1 on Artificial Analysis's Intelligence Index at 57.6 on v4.3.2 and Claude Sonnet 5.5 (September 28) is second at 56.0, ahead of Claude Fable 5.1 at 53.4, GPT-6 Astra at 52.7 and Google's new Gemini 4 Argon at 52.6. Argon, announced September 30, also leads Arena's text and creative-writing boards, but only vetted cyber defenders can use it so far. For coding, Opus 5.5 holds the crown on Arena WebDev and LiveBench Agentic Coding, while Sonnet 5.5 leads Terminal-Bench 4.0 and LiveBench Coding at $2 / $10 and is the value pick. For daily chat, GPT-5.6 is ChatGPT's default and the assistant most people can actually open. For writing, Claude Fable 5.1 has the best combined placing across the three prose boards. For accuracy, Claude Fable 5.1 still holds the highest factual accuracy Artificial Analysis has measured at 67%, with GPT-6.1 Sol the value pick. For hard math and reasoning, GPT-6 Astra leads LiveBench Reasoning and ARC-AGI-2, with GPT-6.1 Sol less than a point behind on both. For images, ChatGPT Images 2.5 leads all four boards we track, and for video MiniMax H3 leads Arena's image-to-video board and is in the top four on three of the four video boards we track. For agents, Gemini Spark is the 24/7 cloud agent and Claude Cowork the desktop one. Among open weights, Xiaomi's MiMo-V2.6-Pro scores 46.3 under plain MIT at $0.13 an index task. GPT-6 Luna is the cheapest model per Intelligence Index task on this page at $0.07, and Claude Sonnet 5.5 is free on claude.ai.
Is GPT-6 out, and is it the best model now?
GPT-6 is out. OpenAI launched it as GPT-6 Astra on September 3, 2026, which settles the year-long question of whether Astra would ship as GPT-6 or as another GPT-5 point release. It is not the top model on the independent boards, and it slipped further on September 22. Artificial Analysis scores it 52.7 on Intelligence Index v4.3.2, 5.7 points clear of GPT-5.6 Sol but 0.7 behind Claude Fable 5.1 at 53.4, 3.3 behind Claude Sonnet 5.5 at 56.0 and 4.9 behind Claude Opus 5.5 at 57.6, while charging $10 / $50 per 1M tokens against Sol's $4 / $20 and Opus 5.5's $4 / $20. Coding was its strongest claim and it is now a split one: Astra took Artificial Analysis's Terminal-Bench 4.0 and held it until Opus 5.5 drew level at 59.6% and Claude Sonnet 5.5 went past at 63.6%, and Arena has since put Opus 5.5 above it on WebDev, leaving Astra only the image-to-WebDev board. It also roughly halves hallucination on AA-Omniscience, from Sol's 92% to 51% at max effort, and it still wins Zapier's AutomationBench and Terminal-Bench-Science 0.1 against Opus 5.5. You need a paid plan: Astra is the first model OpenAI has designated Critical for cybersecurity, so approved defenders in its Daybreak program got it first, with ChatGPT Business and Pro following on September 4 and Plus a few hours later, and nothing has been announced for the free tier. Our full GPT-6 Astra breakdown covers the benchmarks and the caveats. The family grew on September 22 with GPT-6 Sol at $2 / $10 and GPT-6 Luna at $0.10 / $0.50, which Artificial Analysis scores 47.6 and 38.1 against Astra's 52.7. Both are in ChatGPT Work, Codex and the API, and OpenAI says neither is in Chat yet; our GPT-6 Sol and Luna breakdown has the rest. OpenAI added GPT-6.1 Sol on September 29 at the same $2 / $10 as GPT-6 Sol, and Artificial Analysis scores it 51.8, 0.9 behind Astra at a fifth of the token price. There is no GPT-6.1 Astra: the Wall Street Journal reported that OpenAI cancelled its October release after internal safety tests.
Is Gemini 4 out?
Partly. Google announced Gemini 4 Argon, the first Gemini 4 model, on September 30, 2026, but it went only to vetted cyber defenders in Google's Fairwind Program. Google says paid API customers and Google AI Ultra subscribers come next, followed by developers, enterprises and consumers "as soon as possible", and it has given no date or said anything about the free Gemini app. On the independent boards it is a genuine frontier model: Artificial Analysis scores it 52.6 on Intelligence Index v4.3.2, level with GPT-6 Astra, with the lowest hallucination rate it has measured among leading models, and Arena ranks it first on text and creative writing. It is weaker at coding than the leaders, ninth on Arena WebDev and 57.1% on Terminal-Bench 4.0. The introductory API price is $2 / $10 per 1M tokens, rising to $4 / $20. Gemini 3.5 Pro, which Google announced in May, has still not shipped. Full detail in our Gemini 4 Argon breakdown.
What is Claude Opus 5?
Claude Opus 5 is Anthropic's July 24, 2026 flagship, and it was the #1 model on Artificial Analysis until Claude Fable 5.1 arrived on September 1. Claude Opus 5.5 replaced it on September 22, and Anthropic now lists Opus 5 as a legacy model. On Intelligence Index v4.3.2 it scores 50.8, behind Opus 5.5 at 57.6, Claude Sonnet 5.5 at 56.0, Claude Fable 5.1 at 53.4, GPT-6 Astra at 52.7, Gemini 4 Argon at 52.6 and GPT-6.1 Sol at 51.8, and it is fourth among distinct models on both of Artificial Analysis's agentic boards, with 1662 on AA-Briefcase v1.1 and 1724 on GDPval-AA v2.1. API pricing is $5 / $25 per 1M tokens, the same as Opus 4.8 and more than Opus 5.5's $4 / $20, with a 1M-token context window and a May 2026 knowledge cutoff. ARC Prize independently confirmed Anthropic's launch claim on ARC-AGI-3, where Opus 5 scored 30% against 8% for the next-best model at the time, though GPT-6 Astra has since passed it. Arena places it seventh on WebDev, third on image-to-WebDev, first on Document and thirteenth on text. One thing to keep in mind: Artificial Analysis measures its hallucination rate at 61%, which is why it holds no accuracy crown on this page.
What is Claude Fable 5.1?
Claude Fable 5.1 is Anthropic's September 1, 2026 release, and it was the highest-scoring model Artificial Analysis had measured until Claude Opus 5.5 arrived on September 22. It now scores 53.4 on Intelligence Index v4.3.2 at max effort, third among distinct models behind Opus 5.5 at 57.6 and Claude Sonnet 5.5 at 56.0. It still holds the highest raw factual accuracy on AA-Omniscience at 67%, which keeps this page's accuracy crown, and the best combined placing on the three prose boards, which keeps the writing crown. It shipped alongside Claude Mythos 5.1, which is the same model with the biology and cybersecurity safeguards relaxed, available by invitation only and with no published eligibility criteria. Pricing is $10 / $50 per 1M tokens, unchanged from Fable 5, but cache reads fall 75% to $0.25, on a 1M-token context window with a June 2026 knowledge cutoff. It is included in Claude Max and Team Premium at roughly 50% of weekly limits and reachable from Pro through usage credits. For most workloads Claude Opus 5.5 is now the better buy, because it scores higher at less than half the token price. Our full Claude Fable 5.1 and Mythos 5.1 breakdown covers the system card and the safeguard split.
Is Claude Fable 5 back?
Yes. Anthropic redeployed Claude Fable 5 on July 1, 2026 after the US government lifted the export-control restriction it had imposed on June 12. It is available again on the Claude API, Claude.ai, Claude Code, and Claude Cowork. Anthropic made Fable 5 permanent in the paid plans on July 20, 2026. Max and Team Premium include it at roughly 50% of regular usage limits; Pro and Team Standard reach it through usage credits with a one-time $100 starting credit. API pricing is $10 / $50 per million tokens. Fable 5 is a Mythos-class model built for long-horizon agentic work with a 1M-token context, and it sits sixth on Arena's image-to-WebDev board (1,622.8) and eighteenth on WebDev, well behind Claude Opus 5.5 for coding.
What is GPT-5.6 and can I use it?
Yes, as of July 9, 2026. GPT-5.6 is OpenAI's next-generation model family, and after a two-week gated preview that began June 26 behind a US-government safety review, it reached general availability on July 9. There are three tiers, least to most capable: Luna, the fast, cheapest tier ($0.20 / $1.20 per 1M tokens); Terra, a balanced everyday model OpenAI says matches GPT-5.5 ($2 / $12); and Sol, the flagship tuned for biology, chemistry, and cybersecurity ($4 / $20). OpenAI cut Luna by 80% and Terra by 20% on July 30, 2026 and left Sol alone, then cut Sol by more than 20% on August 21, 2026 as a promotional rate for about three months. No ChatGPT subscription price has changed. It is now live across ChatGPT, Codex, and the API as OpenAI's default model. One caveat: OpenAI's system card and the external evaluator METR flagged elevated "scheming" behaviour in Sol, so treat it carefully for high-stakes factual work until independent results settle. One naming point since September 22: OpenAI reused Sol and Luna for the GPT-6 generation, so an unqualified Sol or Luna is now ambiguous. GPT-6 Sol is $2 / $10 and GPT-6 Luna $0.10 / $0.50, against $4 / $20 and $0.20 / $1.20 for the 5.6 tiers, which OpenAI still sells.
Is Grok 4.7 out yet?
Yes. SpaceXAI released Grok 4.7 on September 21, 2026, replacing Grok 4.6 as the flagship. Unlike 4.6, which was a post-training refresh of the same base, 4.7 uses a new and larger base model, trained with a longer reinforcement learning run weighted toward jobs that take many hours; SpaceXAI publishes no parameter count. Artificial Analysis scores it 46.3 on Intelligence Index v4.3.2 at high effort against Grok 4.6's 44.3, and it places above GPT-6 Astra on both agentic boards, 1715 to 1542 on GDPval-AA v2.1 and 1644 to 1569 on AA-Briefcase v1.1. Pricing is unchanged at $2 / $6 per 1M tokens on a 500K-token context window, though prompts of 200,000 tokens and above re-bill the whole request at $4 / $12, and a fast variant runs at twice the output speed for twice the price. The cost per Intelligence Index task did move, from $1.86 to $2.73, because 4.7 emits roughly twice the output tokens. It is live in the Grok API, Cursor, Grok Build, third-party coding harnesses and cloud routers. SpaceXAI's announcement says nothing about the consumer Grok app, which still serves Grok 4.6 at $30/month. Grok is also our creativity pick, on product grounds rather than prose quality; for the best-written output, Claude Fable 5.1 holds our writing crown. Full numbers are in our Grok 4.7 breakdown.
Which AI is the best for coding?
Claude Opus 5.5 is the best for coding, and its cheaper sibling is close behind. Opus 5.5, released September 22, leads Arena's WebDev board at 1815, ahead of GPT-6 Astra at 1788 and Claude Sonnet 5.5 at 1786, and scores 71.7 on LiveBench Agentic Coding, second only to DeepSeek V4.1-Flash and 15 points above Claude Sonnet 5.5, at $4 / $20 per 1M tokens. Sonnet 5.5, released September 28 at $2 / $10, leads Artificial Analysis's Terminal-Bench 4.0 at 63.6% against Opus 5.5's 59.6% and LiveBench Coding at 91.4 against 89.3, which makes it the value pick as long as you keep it below max effort, where an index task costs $7.67 against Opus 5.5's $5.98. GPT-6 Astra is third and keeps Arena's image-to-WebDev board at 1733. Google's new Gemini 4 Argon leads its own DeepSWE v1.1 table at 77.9% but scores 57.1% on Terminal-Bench 4.0 and sits ninth on Arena WebDev, and it is not yet available outside Google's Fairwind program. Among open weights, GLM-5.3 is strongest on Terminal-Bench 4.0 at 41.9% and Kimi K3 is the highest-placed open model on Arena WebDev at thirteenth, while GLM 5.3 Flash is the one most teams can host, at $0.25 an index task under plain MIT. Artificial Analysis has retired both its Coding Index and its Coding Agent Index.
Which AI is the best for writing?
Claude Fable 5.1 is the best for writing, because it has the best combined placing across the three prose boards: second on EQ-Bench Creative Writing v3, third on LiveBench Language and tenth on Arena's creative-writing board. Google's Gemini 4 Argon now leads Arena's creative-writing board at 1518.8, but the other two boards have not rated it and it is not yet available outside Google's Fairwind program. Claude Opus 5.5 is second there at 1516 and the pick if you trust Arena's human votes most, though it is only seventh on EQ-Bench and eighth on LiveBench Language. Claude Fable 5 tops LiveBench Language at 90.7 and is third on Arena's creative-writing board, but eleventh on EQ-Bench. Fable 5.1 and Fable 5 cost $10 / $50 per 1M tokens. Claude Sonnet 5.5 is the value pick and what most people should actually use, since it is free on claude.ai and $2 / $10 in the API, though Arena's creative-writing board puts it only 42nd and EQ-Bench has not rated it yet. For professional deliverables, Claude Opus 5.5 leads GDPval-AA v2.1 at 1866 with Sonnet 5.5 27 points behind at 1839. Kimi K3 is fifth on EQ-Bench at 2082.3 if you want distinctive fiction, GPT-5.5 is the alternative for fact-anchored business writing, and Gemini 3.8 Flash is the price-performance pick for bulk content.
What is the best open-weight AI model in 2026?
Xiaomi's MiMo-V2.6-Pro, published on September 21, 2026 under plain MIT. Artificial Analysis measures it at 46.3 on Intelligence Index v4.3.2, the highest open score it has ever published, above GLM-5.3 at 44.8 and Kimi K3 at 43.6, and it runs an index task for $0.13, cheaper than any other open model on this page. It is a 1.02-trillion-parameter mixture of experts with 42 billion active, a 1M-token context and the weights on Hugging Face from day one. Two caveats: Arena's human voters rate it only 24th on WebDev and 26th on text, and it has barely ten days of independent track record. Kimi K3 remains the highest-placed open model on Arena WebDev, thirteenth, and is now the highest open model on the Agent Arena too, fourteenth, and GLM-5.3 is still the strongest open model on Artificial Analysis's Terminal-Bench 4.0 at 41.9%, on a bespoke Z.ai licence rather than MIT. For teams that actually want to host their own, GLM 5.3 Flash (Z.ai, MIT) remains the practical recommendation at 41.8 on the index and 320B total with 18B active, because a 1.02-trillion-parameter model needs a cluster and the Kimi K3 download is 96 shards and about 1.56 TB.
What is the cheapest frontier-class AI model?
On API pricing per million tokens, GPT-6 Luna is the cheapest closed frontier-class model at $0.10 / $0.50, after OpenAI halved it on September 22, 2026, and Artificial Analysis scores it 38.1 on its v4.3.2 index against Gemini 3.6 Flash's 34.0. Among open weights the pick is still GLM 5.3 Flash, our price-performance pick since August, now at its plain $0.15 / $0.50 list since the launch promotion expired on September 9, scoring 41.8 under a plain MIT licence you can self-host. DeepSeek held that pick until it raised rates roughly fourfold on August 16, and it came close to taking it back on September 10, when V4.1-Flash replaced V4-Flash at $0.15 / $0.60 off-peak and $0.30 / $1.20 at peak: off-peak the two now tie on input, with GLM cheaper on output. Artificial Analysis still separates them on cost per index task, $0.25 for GLM against $0.27 for DeepSeek. Google's cheap tier improved again on September 2, with Gemini 3.8 Flash scoring 40.9 at the same introductory $0.75 / $3.75, though that rate doubles on January 1, 2027. MiniMax M3 lists at $0.30 per million input tokens on a permanent 50% discount, with LongCat-2.0 as a strong MIT alternative. Measured per Intelligence Index task rather than per token, the cheapest model on this page is GPT-6 Luna at $0.07, ahead of GPT-5.6 Luna at $0.18 and the open-weight MiMo-V2.6-Pro at $0.13.
Which AI models are free?
ChatGPT Free now defaults to GPT-5.6 (with GPT-5.5 still available) under usage limits. Free and Go users can also reach GPT-6 Luna in the ChatGPT desktop app since September 22. Gemini Free runs Gemini Flash-Lite in the Gemini app from October 9 (3.6 Flash before), and Gemini 3.6 Flash stays free in Google AI Studio. Claude Free runs Claude Sonnet 5.5, released September 28, with usage limits. DeepSeek Chat runs DeepSeek V4 free on the DeepSeek website. Grok has a limited free consumer plan (X Premium is a paid add-on). Qwen 3.5, Qwen3.8-Flash-Next, NVIDIA Nemotron 3 Ultra, MiniMax M3, LongCat-2.0, Kimi K2.6, DeepSeek V4, GLM-5.2, GLM 5.3, GLM 5.3 Flash, MiMo-V2.6-Pro and Tencent's Hy4 Preview are open-weight and free to self-host, though the licences differ and only some of them are plain MIT or Apache. Qwen 3.7 Max is API-only with no consumer chat front-end, but Alibaba now includes 200 free model requests per day. Kimi has a free basic tier in its app, with heavier agentic use metered.
What is Gemini Spark and which plan do you need?
Gemini Spark is Google's first 24/7 cloud-resident AI agent, launched at Google I/O on May 19, 2026 and no longer an Ultra exclusive: since July 30, 2026 it also reaches the $19.99/month Google AI Pro tier, in the US and more than 160 further countries, while Google AI Ultra at $99.99 or $199.99/month carries the wider global rollout. Google AI Plus, the free tier and work or school accounts do not include it. Spark is built on Gemini base models with Google's Antigravity harness on a Google Cloud VM, integrates with Gmail, Google Docs, and other Google Workspace apps, and can interact with Chrome and Android's Halo system on the device side. It is worth the spend for users who have repeatable long-running workflows (inbox triage, research roll-ups, scheduled tasks). For one-off tasks, Claude Cowork at $20/month covers most desktop-agent needs.
What is Fello AI?
Fello AI is an AI chatbot for Mac, iPhone, and iPad that lets you use all top AI models like ChatGPT, Claude, Gemini, Grok, and DeepSeek in one app, with models updated regularly so you always have the latest. It is $9.99/month with a 4.7-star rating across 27,000+ reviews.
How often do you update this page?
We update this page at least monthly and within 24-48 hours of any major model launch.
Related Articles
Claude vs ChatGPT: Which AI Is Actually Better in 2026?

Claude hit #1 on the App Store in early 2026, pushing ChatGPT out of the top spot for the first time. The catalyst was Anthropic publicly refusing the Pentagon's demand to deploy its models for autonomous weapons.

Read more →
Grok vs ChatGPT: Which AI Chatbot Is Actually Better in 2026?

Neither chatbot serves its lab's newest model yet: ChatGPT chats still run on GPT-5.6 Sol and the Grok app on Grok 4.6, while GPT-6.1 Sol and Grok 4.7 (September 21, $2 / $6) sit in ChatGPT Work, Codex and SpaceXAI's developer tools.

Read more →
ChatGPT vs Gemini in 2026: Which AI Should You Actually Use?

ChatGPT still answers chats with GPT-5.6 Sol and the Gemini app still serves Gemini 3.1 Pro, while GPT-6 and Gemini 4 Argon sit behind Work, Codex and a Fairwind-only rollout. Head-to-head across writing, coding, images, and price.

Read more →
When to Use Which AI: Pick the Right Model

There is no single best AI. Match the task to the model that leads it, with a task-by-task table for coding, writing, research and everyday chat.

Read more →
Best AI Agents in 2026: 30 Tools Tested

Thirty AI agents compared by what you actually want to do: coding, research, office work and automation, with honest pricing and free options.

Read more →
Gemini Nano Banana Pro vs GPT-Image-1.5: Ultimate Comparison

Our original December 2025 hands-on comparison of what were then the two leading image generation models, with real output samples side by side. Both have since been succeeded.

Read more →

Try every model.
One beautiful app.

Every model from this guide in one native app: ChatGPT, Claude, Gemini, Grok, and DeepSeek. Free to start.

Fello AI running on Mac, iPad, and iPhone

4.7 rating·27,000+ reviews·Free to start