Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

Model Evaluations and the Messy Business of Measuring Danger

A dangerous capability evaluation asks one specific question: can this model do the thing we are worried about? The specificity is what makes it useful and what makes it hard.

Scientific measurement equipment in a laboratory — photo by Belikov Maxim, licensed under CC BY 4.0 via Wikimedia Commons.

A dangerous capability evaluation attempts to answer a specific question: can this model do the thing we are worried about? The specificity is what makes it useful and what makes it hard. A vague question returns a vague answer, and a vague answer cannot be written into a rule.

Four things people test for

  • Cyber: can the model find or exploit vulnerabilities in code or systems?
  • Biological: can it provide meaningful uplift to someone attempting to cause harm?
  • Autonomy: can it pursue a goal across many steps, including evading oversight?
  • Persuasion and deception: can it mislead a person in ways they cannot detect?

A failing model may just be a badly prompted one

Elicitation is the first problem. A model that fails a test may simply not have been prompted well. A model that passes may have been lucky. Evaluators spend much of their time trying to extract a model’s true capability, which means the measurement depends on the skill of the measurer as much as on the model.

Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.

The second problem is baseline. Uplift over what? A capable person with search access already has substantial capability, so an evaluation has to compare against a realistic control, and constructing that control is methodologically demanding enough that it rarely gets done properly.

The third is generalisation. A model that fails on one version of a task may succeed on a paraphrase, or on a task in a different language.

A passing score is evidence, not a certificate

Because no evaluation is definitive, the useful regime is one that requires disclosure of method and results, mandates external review, and treats a passing score as one piece of evidence rather than a certificate. Regulators who treat evaluations as proof will be disappointed, and the disappointment will arrive on a schedule of its own.

Red teams find the things nobody wrote a checklist for

Structured evaluation is supplemented by adversarial testing, in which people attempt to make a system misbehave. Red-teaming finds problems that no checklist anticipates, and it produces findings that are difficult to generalise.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

The methodological difficulty is that a successful attack proves vulnerability while a failed attempt proves nothing. A team that finds no way to bypass a safeguard cannot conclude that none exists, which makes the negative result uninformative and the resource allocation hard to justify to whoever holds the budget.

Publishing the method helps the attackers too

When evaluators find a serious issue, the decision about what to publish is genuinely difficult. Publishing the method helps others test their own systems; not publishing prevents the method from being used against systems that remain vulnerable.

Current practice varies and is largely discretionary. The absence of a standard disclosure protocol is one of the more consequential gaps in the field, because it means the public learns about failures at the discretion of the organisations that experienced them.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.