Every AI model launch now arrives with a wall of numbers, and almost nobody explains them. AI benchmarks are standardised tests, and each one measures something far narrower than the headline implies. When Anthropic, OpenAI or Google says a model scores 59 on one test and 80 on another, those two numbers come from different question sets, different scoring rules and sometimes different versions of the same benchmark.

The gap between what a benchmark measures and what people think it measures is where most confusion lives. This guide walks through the tests you actually see quoted, MMLU, GPQA, Humanity's Last Exam, SWE-bench, Terminal-Bench, LiveBench and FrontierMath, plus the two leaderboards that aggregate them. It ends with the part nobody writes down: how much any of it should change the model you open on your Mac tomorrow morning.

The Key Takeaways

  • A benchmark is a fixed test, not a verdict: the same questions, asked the same way, scored the same way, measuring one narrow skill.
  • The same benchmark gives different numbers: on Humanity's Last Exam, Claude Fable 5.1 scores 59.1%, 58.7% and 55.9% on Artificial Analysis depending only on the reasoning effort setting.
  • The headline index does not contain the famous tests: the Artificial Analysis Intelligence Index v4.2 runs ten evaluations, and neither MMLU nor GPQA is one of them.
  • Leaderboards measure different things: Arena ranks by blind human preference, the Intelligence Index aggregates capability evals. A model can lead one and not the other.
  • Version drift is real: Terminal-Bench is on v4.0 while the index that quotes it runs v2.1, so two correct sources can disagree.

What Is an AI Benchmark?

Do editor

Todos os modelos de IA numa só aplicação

Fello AI reúne GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 e mais numa só aplicação nativa para Mac e iPhone.

Descarregue já!

An AI benchmark is a fixed set of questions or tasks given to every model in the same way and scored the same way, so results can be compared. That is the whole idea.

A benchmark measures one narrow skill, graduate-level chemistry, fixing real bugs in real code, or which of two answers a human prefers, and no single one describes how good a model is. It is also why a single run proves little, given how much one model's answer varies between attempts.

Two things share the word and are not this. Hardware benchmarking measures how fast a chip or a video app runs, which is a throughput question about your machine, not about a model's ability. And marketing teams now talk about benchmarking brand visibility inside AI answers, which is a search-analytics question. This guide is about model evaluations only.

Why the Numbers Keep Changing

Benchmarks decay. Once models reach the ceiling, a test stops separating them, so a harder replacement appears and the leaderboards migrate. That is why the tests quoted in 2026 are not the tests quoted in 2024, and why a guide written six months ago will teach you the wrong scoreboard. It is also why our ranking of the best AI models is rebuilt monthly rather than left standing.

The Knowledge Tests: MMLU, MMLU-Pro and GPQA

These are the multiple-choice exams. They ask whether a model knows things and can reason over what it knows, and they are the oldest family still in circulation.

MMLU

Measuring Massive Multitask Language Understanding covers 57 tasks, from elementary mathematics and US history to computer science and law. For years it was the headline number on every model card. It is now largely finished as a discriminator: frontier models cluster near the top, so a two-point difference tells you almost nothing.

MMLU-Pro

The successor tightens the screws in two ways. It expands the answer options from four to ten, which cuts the reward for lucky guessing, and it selects harder questions. The MMLU-Pro paper reports a 16% to 33% drop in accuracy compared with MMLU, and a useful side effect: sensitivity to prompt wording fell from 4 to 5% down to about 2%. A more stable test is a better test, even before it is a harder one.

GPQA and GPQA Diamond

GPQA is 448 multiple-choice questions written by domain experts in biology, physics and chemistry, and designed to be Google-proof. The headline evidence is the human gap: people with or pursuing PhDs in the matching field reach 65% accuracy, while highly skilled non-experts reach only 34% despite spending on average more than 30 minutes with unrestricted web access. Most quoted scores use the Diamond subset, which Epoch AI describes as the 198 questions the expert validators got right and most non-experts got wrong. When you see GPQA in a launch chart, assume Diamond unless told otherwise.

The Hard Reasoning Tests: Humanity's Last Exam and FrontierMath

When the multiple-choice exams saturated, the field built tests deliberately positioned beyond current models. These are the ones where a 60% looks unimpressive and is actually extraordinary.

Humanity's Last Exam

Built by the Center for AI Safety and Scale AI with nearly 1,000 expert contributors from more than 500 institutions across 50 countries, Humanity's Last Exam is 2,500 questions across more than a hundred subjects. It is the closest thing the field has to a general frontier exam, and it is the single benchmark most likely to be quoted at you with a number that does not match another number you saw yesterday. That is not sloppiness on anyone's part, and the reason is worth its own section below.

FrontierMath

Epoch AI's FrontierMath is several hundred unpublished mathematics problems running from undergraduate difficulty up to genuine research level, plus a set of open problems with computationally verifiable solutions and a Lean-formalised Erdős collection. Unpublished is the whole design: problems that never appear online cannot leak into a training set, which makes it one of the few tests where a high score is hard to explain away.

The Coding and Agent Tests: SWE-bench, Terminal-Bench and LiveBench

This is where the action moved in 2026. Answering a question is not the same as doing a job, and these tests grade the job.

SWE-bench

SWE-bench is 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. The model gets a codebase and a bug report and has to produce a change that makes the repository's own tests pass. You will also see SWE-bench Verified, which the project describes as a subset of 500 problems that real software engineers have confirmed are solvable, and SWE-bench Pro. Three names, three different numbers, one family.

Terminal-Bench

Hosted by Stanford, Harbor and the Laude Institute, Terminal-Bench scores agents on real terminal work and reports a resolution rate alongside cost and token usage. It is also the cleanest example of version drift in the whole field, which the leaderboard section returns to.

LiveBench

LiveBench exists to attack one specific problem: test set contamination, where benchmark questions end up inside a newer model's training data. It draws questions from recently released competitions, arXiv papers, news articles and datasets, adds and updates them monthly, and scores against objective ground truth rather than an LLM judge. If you only trust one number on a model that launched last week, a contamination-limited one is the better bet.

The Two Leaderboards, and What Each One Is Really Ranking

Almost every score you see quoted secondhand comes from one of two places, and they are measuring different things.

Arena, Formerly LMArena

Arena ranks models by blind human votes. You send one prompt, two anonymous models answer, you pick the better reply, and identities are revealed afterwards. Since December 2023 the ratings are not classic Elo. As the team wrote when they made the switch, the core difference is the assumption that a player's performance does not change. That lets them fit all votes at once by maximum likelihood, and produce bootstrap confidence intervals that actually reflect the uncertainty. Worth knowing: lmarena.ai now redirects to arena.ai and the site brands itself Arena.

The scale is easy to misread. Under the rating formula a 100-point gap implies roughly a 64% win rate, not dominance, and frontier models routinely sit within 20 to 40 points of each other. Higher also means preferred, not correct. A confident, well formatted, slightly long answer wins blind votes it does not always deserve.

The Artificial Analysis Intelligence Index

The Intelligence Index is a composite. The current v4.2 methodology runs ten evaluations: AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR v1.1, AA-Omniscience, Humanity's Last Exam, GDP.pdf and CritPt. It weights agents at 30%, coding at 20%, scientific reasoning at 20% and general at 30%, scores with pass@1, and puts a 95% confidence interval of under one percentage point on the result.

Two details in that list deserve attention. Neither MMLU nor GPQA is in it, so the tests most readers ask about by name no longer feed the index number they are reading. And it runs Terminal-Bench v2.1 while Terminal-Bench itself is on v4.0. Both are correct. They are simply not the same measurement, which is exactly how two honest sources end up disagreeing.

Why the Same AI Benchmark Gives Different Numbers

This is the part that turns benchmark literacy from trivia into something useful. Take Humanity's Last Exam on a single day. Read on 7 September 2026, its own leaderboard puts Gemini 3 Pro on top at 38.3%, with GPT-5 at 25.3% and Grok 4 at 24.5%. The same day, the Artificial Analysis board for the same benchmark shows Claude Fable 5.1 at 59.1%. Neither is wrong.

The Four Reasons a Score Moves

First, the question set. Artificial Analysis states plainly that it evaluates models on 2,158 text-only questions from the exam rather than all 2,500, because not every model is multimodal. Second, the settings. On that same board, Claude Fable 5.1 scores 59.1% at max effort, 58.7% at extra-high and 55.9% at high, a three-point spread from one identical model on one identical test. Third, tool use, which some harnesses allow and others forbid. Fourth, the calendar: a leaderboard that has not re-run recent models tells you about last quarter.

How to Read a Quoted Score

The practical rule is short. A benchmark number without its source, its variant and its settings is decoration. When you see one, ask which board it came from, which subset was used, and whether the model was allowed to think longer or use tools. If a comparison chart cannot answer those, it is marketing. Our own coverage of individual launches, such as the Claude Fable 5.1 benchmark run, names the board and the settings for exactly this reason.

What AI Benchmarks Cannot Tell You

Every serious evaluator says some version of this, and every marketing chart ignores it.

Saturation and Contamination

A saturated benchmark is one everybody passes, which is why MMLU stopped being interesting. Contamination is worse, because it is invisible: if a test leaked into training data, a high score measures memory rather than ability. That is the entire reason LiveBench refreshes monthly and FrontierMath keeps its problems unpublished.

The Gap Between Scores and Daily Use

A model can win the board and annoy you within an hour. Benchmarks do not measure refusal behaviour, tone, response speed, how gracefully a model handles a vague request, or whether it keeps its promises across a long conversation. We wrote about a model that topped the benchmarks and still frustrated its users, and that pattern has repeated with every generation since. If you want the mechanics behind the scores, our explainer on how AI reasoning models actually think covers what the reasoning effort setting is really doing.

Which AI Benchmarks Actually Matter for Your Work

Match the test to the job. Nothing else about this is complicated.

If you mostly do thisWatch thisIgnore this
Everyday chat and writingArena preference ratingsGraduate science scores
Research and technical readingGPQA Diamond, Humanity's Last ExamHuman preference votes
Coding and debuggingSWE-bench Verified, LiveBench codingGeneral knowledge tests
Agents and multi-step tasksTerminal-Bench, the agent-weighted indexSingle-turn exams
Maths-heavy workFrontierMathMMLU
Running a model locallyOpen-weight scores at your parameter sizeFrontier leaderboard positions

One more filter matters as much as any score: what the model costs to reach. A three-point index difference is rarely worth a doubled subscription, which is the calculation in our full AI pricing comparison. And if you are choosing something to run on your own machine, our ranking of the best open source AI models is scored on the same index discussed above.

The Case for Not Chasing the Leader

Leaderboards change monthly, indices get re-versioned, and the model that wins in September is often not the one that wins in November. Committing a year of subscription to whoever topped a board last week is the expensive way to use this information. Keeping several frontier models a keystroke apart, and sending each task to whichever one actually suits it, is how Fello AI users sidestep the question entirely.

The Verdict

Benchmarks are useful and they are not a ranking of intelligence. They are narrow instruments with versions, subsets and settings, and the honest way to read one is to ask what it tested, on which subset, with what effort allowed, and on what date. Do that and the numbers become informative.

Skip it and you are reading a marketing chart with decimal places. The practical move is to pick the one or two tests that match your actual work, treat everything else as background noise, and keep more than one model within reach so a leaderboard shuffle never costs you anything.

FAQ

What is an AI benchmark in simple terms?

It is a fixed set of questions or tasks given to every model in the same way and scored the same way, so the results can be compared. Each benchmark measures one narrow skill, such as graduate science or fixing real code, and no single one describes how good a model is overall.

Why do two leaderboards give the same model different scores?

Because they are not running the same test. Boards differ on which subset of questions they use, whether tools are allowed, how much reasoning effort the model is given, and when the model was last re-run. Artificial Analysis, for example, evaluates Humanity's Last Exam on 2,158 text-only questions rather than the full 2,500.

Are AI benchmarks reliable?

They are reliable about the narrow thing they test and unreliable as a general verdict. Two known failure modes are saturation, where every model passes and the test stops separating them, and contamination, where questions leak into training data and a high score reflects memory rather than ability.

Does a higher Arena rating mean a smarter model?

It means people preferred its answers in blind comparisons, which is not the same as correctness. A 100-point gap implies roughly a 64% win rate, and frontier models often sit 20 to 40 points apart. Longer, confidently formatted answers tend to do well whether or not they are right.

Which benchmark should I look at before choosing a model?

Match it to your work. Arena for everyday chat and writing, GPQA Diamond or Humanity's Last Exam for research, SWE-bench Verified and LiveBench for coding, Terminal-Bench for agent tasks, FrontierMath for mathematics. Then weigh the score against what the model costs you each month.