Percy Liang and the Leaderboards Everyone Argues About
Percy Liang measures what these machines can do. The trouble with any scoreboard, as every gambler knows, is what men will do once a prize is posted beside it.
Every large claim about artificial intelligence rests on a measurement, and the measurement is almost always a benchmark. That is a word I have learned late and come to distrust in the ordinary way — not because the men who make benchmarks are dishonest, but because benchmarks are made by people who would be the first to tell you their limits, and nobody ever quotes those people in a headline. I have watched a great many men win an argument by producing a number, and I have never once seen the number examined on the way to the podium.
Percy Liang has led several of the most widely cited evaluation efforts in the field, including broad assessments that compare models across reasoning, knowledge, coding, and the behaviours that matter when a person actually uses an assistant. Which is to say he has volunteered for the least romantic job in the enterprise — the counting, and then the defending of the counting.
The four ways a scoreboard lies
- Contamination: training data scraped from the web may contain the test set itself.
- Overfitting to the format rather than to the underlying capability.
- Metric substitution: the number goes up while the user’s experience does not improve at all.
- Distributional narrowness: tasks that bear no resemblance to anybody’s actual work.
What a decent measurement looks like
The direction of travel is toward held-out private sets, tasks designed by someone actively trying to defeat the model, human preference evaluations built with careful sampling, and — increasingly — evaluations of behaviour under pressure. What a model does when it is uncertain. What it does when asked for something it ought to refuse. What it does when handed a task it cannot possibly complete. Those are the moments that tell you what you own, and they are precisely the moments a clean demonstration never shows.
Liang has argued that the most useful evaluations are unglamorous ones: they measure specific, narrow behaviours reliably, rather than producing a single headline score that compresses everything into a ranking. I have seen what men do with a single headline score. They print it on a banner, and the banner outlives the measurement by a decade.
The judge is also a shareholder
The awkward structural fact is that the laboratories being evaluated often fund the evaluations, or contribute to them, or employ the people who run them. Independent measurement requires resources that academic groups rarely possess, and asking a man to grade his own examination is a practice I have seen before, in towns with one newspaper. That is a regulatory problem as much as a technical one, and it is why disclosure requirements keep turning up in draft legislation, drafted by people who have been handed too many banners.
The places these machines still fall down
Systematic evaluation has produced a consistent picture of where these systems break: multi-step arithmetic with large numbers, spatial reasoning, tracking entities through a long document, recognising when a premise is false, and holding a consistent position across a long interaction. I recognise that last one. It is a complaint I have lodged against several acquaintances.
Those failures are more informative than any aggregate score, because they point at a mechanism rather than a mood. A system that handles a problem in one phrasing and fails on a synonymous one is not reasoning about the problem. It is matching a pattern, the way a man may answer a question correctly by recognising the sound of it rather than the sense.
The race between the measure and the measured
As soon as a benchmark becomes a target, it begins to decay. Public sets get swept into training data, formats get optimised against, and the scores climb faster than the capability they are meant to count. Every serious evaluator now assumes contamination and designs accordingly, which is progress of a sort — the same progress a banker makes when he assumes every visitor has read the balance sheet.
The practical consequence is that evaluation is a permanent research activity rather than a piece of work one finishes and files. It requires continuous investment, and any organisation that treats it as a project with an end date will find its measurements meaningless within the year. I have watched a great many men build a fine fence and then leave the gate open, and it always ends the same way.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.