Should You Trust AI Benchmarks When Picking an Assistant?

Every model launch comes with a chart. Bars, percentages, a rival or two trailing behind. And every launch, a lot of people make a purchasing decision on the basis of that chart.

The short answer: benchmarks are genuine measurements of things that are probably not your work. They’re useful for tracking whether the field is moving and for filtering out models that are clearly not in the running. They’re close to useless for choosing between two mature assistants for your own use, and knowing why makes you much harder to sell to.

We’re not going to print any scores on this site — partly because they’d be stale within weeks, and partly because reproducing a leaderboard is not the point. The point is how to read one.

What benchmarks are good for

Detecting a real generational jump. When a new model class clears a bar that nothing previously cleared, that’s a signal worth noticing. Big jumps are real.

Filtering the obviously unsuitable. If you’re considering a small model for a demanding job, benchmarks will tell you it’s out of its depth faster than your own testing will. Cheap negative screening.

Tracking the field over time. Watching how quickly open-weight models close the gap on hosted ones is a genuinely useful macro signal, and it’s the one thing benchmarks measure well. See what are open-weight models.

Narrowing a large shortlist. Twelve candidates down to four, cheaply. Then stop.

The four ways they mislead

1. They measure a different task than yours

Most well-known benchmarks are collections of short, self-contained problems with a checkable answer: exam-style questions, isolated coding puzzles, knowledge quizzes. Your work is probably long, contextual, ambiguous, and iterative — a report to restructure, a codebase to change without breaking, a draft to revise toward a voice.

A model can be excellent at short verifiable problems and merely fine at sustained work with your material, or the reverse. The correlation is positive but far weaker than the charts imply. Nothing on any public leaderboard measures “held the thread through my forty-page document” or “stopped adding adverbs when I asked it to.”

2. The gaps are smaller than they look

Two habits inflate them. Truncated axes, where a bar chart starts at 70% and a three-point difference occupies half the image. And absent error bars — many benchmark results have run-to-run variance comparable to the differences being celebrated, so a small lead may not be a lead at all.

The practical filter: treat small differences between mature models as noise unless someone shows you the variance. Vendors rarely do, and that isn’t specific to any one of them.

3. Contamination is real and unfixable

Benchmarks are public. Public text ends up in training data. When a model has seen a test’s questions and answers, its score measures recall rather than capability, and no one can fully prove it didn’t happen.

Labs work hard on this — held-out sets, fresh problems, private evaluations — and it’s a genuine effort, not a sham. But contamination is structurally hard to rule out, which means an old, famous, heavily-cited benchmark is the least trustworthy kind. A brand-new benchmark is more informative than a well-known one, precisely because it’s had less time to leak.

4. Vendors choose which numbers to show

Not fraud — selection. Every lab has a large internal matrix of results and publishes the flattering slice, with the configuration that helped: a particular prompting technique, a particular sampling setting, the specific comparison model that makes the gap widest. Every vendor does this. When you see a chart in a launch post, you’re seeing the best true thing that could be said.

The tell is asymmetry in configuration detail: if the vendor’s model is described with a technique and the comparison isn’t, the comparison probably didn’t get one.

The special case of human preference rankings

Crowd-voted arenas, where people blind-compare two answers and vote, dodge some of these problems — the prompts are real, the judgement is human, and there’s nothing to memorise.

They introduce others. They reward answers that look good to a quick reader: well-formatted, confident, appropriately long. That’s a real quality, but it’s not the same as being correct or being right for a specialist. The voter population also isn’t you — their prompt mix is not your prompt mix, and their taste in prose is not necessarily your taste.

Treat these rankings as a decent proxy for “pleasant general-purpose answers” and a poor proxy for “good at my domain.” Which, notably, is exactly what you’d conclude about any popularity ranking.

What to use instead

The signature move on this site, because it works: build your own tiny benchmark. It takes an afternoon once and pays out for years.

  1. Collect five real tasks from your own recent work, spanning what you actually do. Real inputs — long ones if your work is long, messy ones if your work is messy. Toy prompts flatten every difference you care about.
  2. Write down what a passing answer looks like before you run anything. This is the step people skip, and skipping it turns the whole exercise into rationalisation.
  3. Run all candidates with identical prompts, including any style sample or context you’d normally supply. Coaching one side is how people fake their own results.
  4. Add one hard correction per task. “No — keep my phrasing in paragraph two and cut a third.” How a model handles being told it got it wrong is more decisive in daily use than its first draft, and no public benchmark measures it.
  5. Judge blind. Strip the labels, shuffle, choose. Brand knowledge contaminates preference far more than anyone believes about themselves.
  6. Keep the set. When a new model ships, you have a twenty-minute test instead of a chart to squint at. This is the real payoff.

Five real tasks with a written pass criterion will tell you more about which assistant to use than every leaderboard combined, because it’s the only measurement weighted by your task distribution.

Bottom line

Read benchmarks the way you’d read a manufacturer’s fuel-economy figure: measured honestly under specified conditions, useful for comparing generations, not a prediction of what you’ll get. Use them to shortlist and to notice genuine jumps. Never use them to pick between two mature products for your own work — that’s what your own five tasks are for.

For the full decision procedure, see our framework for choosing an alternative, and is Claude better than ChatGPT for what a head-to-head looks like when nobody’s crowning a winner.