Jan Leike: The Alignment Researcher Who Left
A resignation letter from a senior safety researcher became one of the most consequential documents of the year — and it named a specific number.
Corporate resignations are usually forgettable — a paragraph, a thank-you, a forwarding address, and a box carried out to the car by somebody from facilities. The one posted by a departing alignment researcher at OpenAI was not, because it made a specific and checkable allegation: that safety work had been deprioritised relative to shipping, and that the organisation’s stated commitment to the problem was not matched by its resource allocation.
He then joined a competitor, taking the same research agenda with him. A man who resigns in protest and retires to a farm has made a gesture. A man who resigns and takes the argument across the street has made a move.
What alignment work involves
The term covers a wide range of practical research, and it is not, as the layman might suppose, one man in a room worrying. Making a model’s behaviour match stated intentions. Preventing it from pursuing goals that diverge from the ones it was given. Detecting deceptive behaviour. Building systems that can be inspected and corrected.
- Scalable oversight: supervising systems that may exceed their supervisors.
- Honesty and calibration: making a model’s expressed confidence track its actual reliability.
- Robust refusal: preventing a model from being talked out of its constraints.
- Evaluation: measuring whether any of the above is working.
The resource argument
The core claim in the departure was about allocation — that is, about who gets the money and who gets overruled. A safety team that is large in headcount but smaller than the product organisation, and whose recommendations can be overridden on schedule grounds, is not really a check. Whether that describes any particular company is disputed, but the structural point is hard to refute.
What changed afterwards
Publicly, several labs reorganised their safety functions, added external review, and published more about their processes. Whether those changes were substantive or cosmetic is precisely the kind of question outside researchers cannot easily answer without access.
That is the recurring difficulty of the entire field: the evidence needed to evaluate a safety claim is typically held by the organisation making it.
The measurement difficulty
Alignment research has a verification problem that capability research does not. A better benchmark is easy to identify, because the number goes up and the graph is pretty. A claim that a model is honest, or that it will not pursue an objective it was not given, is much harder to falsify, and a system that appears to pass every test may simply be good at passing tests.
This is why the field has gravitated toward interpretability: if you can inspect the computation rather than the output, a claim becomes checkable. It is also why the most-cited alignment results tend to be negative ones — demonstrations that a particular safeguard can be bypassed.
The organisational bind
Every lab faces the same structural issue: the people best positioned to identify risks to a release schedule are employed by the organisation that wants to keep the schedule. That is true of aviation and pharmaceutical safety as well, and the standard solution is an independent regulator with subpoena power.
Until such a regulator exists with technical capacity, safety teams will remain internal advisory functions whose influence depends on the personal standing of the people in them. That is a thin rope to hang a technology on, because standing walks out the door with the man.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.