The Pre-Deployment Testing Debate Nobody Has Settled
The aviation analogy gets reached for first and holds up least well. The fight over mandatory pre-release testing is really a fight about staffing.
The analogy politicians reach for most often is aviation: aircraft must be certified before they carry passengers, and no manufacturer gets to decide for itself whether its design is safe. Why models should be different is the obvious next question.
Airframes have fixed physics. Models do not.
An aircraft is a fixed design with predictable physics. A model is a distribution of behaviours that shifts with the prompt, the fine-tune and the deployment context. Certifying it is less like certifying an airframe and more like certifying a programming language.
There is also the matter of pace. Certification cycles measured in months are incompatible with release cycles measured in weeks, which is why industry proposals favour continuous evaluation over one-time approval.
Four proposals, none of them agreed
- Mandatory pre-release testing with a public authority holding veto power.
- Voluntary submission with published results and no veto.
- Threshold-triggered testing: only systems above a compute or capability line need approval.
- Liability-based regimes that leave testing to the developer but impose consequences for harm.
A few thousand people, mostly inside the labs
Every version of mandatory testing requires a body capable of doing it, and the number of people worldwide who can competently evaluate a frontier model is in the low thousands at most. Most of them work for the companies being tested.
That scarcity, more than any philosophical disagreement, is why the debate has not resolved. A regime that cannot be staffed is not a regime. Building the capability is the unglamorous precondition that everyone acknowledges and few are funding.
What aviation has that AI does not
The comparison to aviation is usually made badly. Aircraft certification involves a small number of manufacturers, a stable design, an established regulator and a clear definition of failure. None of those conditions hold here.
A better analogy may be pharmaceutical approval, where a product is tested against a control, monitored after release, and subject to withdrawal if problems emerge. That regime also took decades to build, and it is considerably slower than the pace of AI releases.
Voluntary access agreements, and what they are worth
In the absence of a mandatory regime, several governments have established voluntary pre-deployment access agreements in which labs share models with public bodies before release. These give officials visibility without granting veto power.
The arrangement has real value as a capability-building exercise and limited value as a safety mechanism. Its defenders describe it as a step toward something with teeth, and its critics describe it as a way to appear to regulate without doing so.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.