Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

Benchmarks Are Broken. Here Is What Should Replace Them

A field that cannot measure what it builds will keep believing its own marketing.

Charts and statistical analysis — photo by Pageview Analysis; I made the screenshot in Firefox, licensed under CC0 via Wikimedia Commons.

Public benchmarks have become the field’s primary scoreboard and its most reliable source of self-deception. The problem is not that they measure the wrong things entirely. It is that they are the things easiest to measure, and the field has optimised against them.

Contaminated, saturated, narrow, disconnected

  • Contamination: test items appear in training data, sometimes verbatim.
  • Saturation: once scores approach the ceiling, differences become noise.
  • Narrowness: a score measures one phrasing of one task under one condition.
  • Disconnection: no correlation with whether a practitioner finds the system useful.

Held-out sets, real tasks, reported agreement

The alternative is more work per measurement and less comparability across systems, which is exactly why it does not happen. It looks like this: private held-out sets that are regenerated regularly, tasks drawn from the actual work the system will do, human evaluation with careful sampling and reported agreement, and explicit measurement of failure modes rather than aggregate scores.

Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.

It also means publishing negative results. A regime that only reports what a system does well is advertising, not evaluation.

Measurement capability is infrastructure with no funder

Independent evaluators need resources and access that only governments or large philanthropic funders can currently provide. That is a policy recommendation dressed as a methodological one, and it is the crux of the issue: measurement capability is infrastructure, and nobody is funding it at the scale the stakes imply.

Buyers want a ranking that this approach will not give them

The alternative to benchmarks is not no measurement; it is measurement that is not comparable across systems. That is a genuine cost, because buyers want to choose between options, and regulators want to set thresholds.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

The compromise most practitioners adopt is layered: broad public benchmarks for rough screening, private domain-specific suites for decisions. The public numbers get attention and the private ones get used.

A favourable result is a press release

Companies publish benchmark results selectively. A favourable result is a press release; an unfavourable one is not mentioned. Across an entire industry this produces a systematically misleading public picture built entirely from true statements.

Requiring pre-registration of evaluation plans, as clinical trials do, would be a heavy-handed solution to a real problem. It is also the only mechanism anyone has devised that reliably prevents selective reporting.

None of that excuses selective disclosure. The correct standard is not that published evaluations must be unfavourable, but that the set of evaluations published should not be chosen after the results are known. Very few organisations currently meet that standard, and it has not become a condition of doing business.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.