Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

Chris Olah and the Effort to Read AI’s Mind

A model has billions of parameters and no comments. Chris Olah has spent his career trying to read it anyway — and it is the only route to verifying anything a model says about itself.

An abstract visualisation of a neural network — photo by mikemacmarketing / original posted on flickr Liam Huang / clipped and posted on flickr, licensed under CC BY 2.0 via Wikimedia Commons.

The word interpretability describes an attempt to answer a deceptively simple question: what is going on inside a trained neural network? A model has billions of parameters and no comments. Its behaviour can be measured; its reasoning cannot be read. I have spent a good part of a long life reading men who had every intention of being understood, and I can report that this is frequently harder than it sounds. A mind with no comments at all is a new order of difficulty.

Chris Olah has been one of the field’s most influential figures in this area, first at a research organisation focused on visual explanations and later leading an interpretability team at Anthropic. I have not met the man, and I will not pretend otherwise, which is the only honest way to open a column of this kind.

Take the machine apart and label the pieces

The most publicised line of work in this area attempts to decompose the activations of a network into features that correspond to interpretable concepts — a direction, a colour, a syntactic role, a named entity — and then to trace how those features combine. If it works, the result is something closer to a circuit diagram than a black box. I once watched a man take a steamboat’s engine apart on a Sunday afternoon and lay every part on the levee in the order he removed it. By evening the boat ran. He could not have told you why a single one of those parts was shaped the way it was, and the boat did not care.

  • Identifying features that a model uses internally and can be steered by.
  • Mapping how features combine across layers to produce behaviour.
  • Testing whether interventions on those features change outputs in predicted ways.
Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.

Why the work goes by the inch

The obstacle is scale and ambiguity. A model may represent a concept across many dimensions, in ways that are not obviously interpretable to a human observer, and any given behaviour may involve thousands of interacting features. Progress is measured in individual mechanisms rather than in whole systems. That is a poor way to write a press release, and a sound way to learn a river. Every man who ever learned the Mississippi did it one bend at a time, and cursed the mapmakers who drew it as a line.

The part that pays for the rest of it

The reason this work matters beyond academic interest is that it is the only route to verification. Regulations that require a company to demonstrate that its model does not do something are unenforceable if nobody can inspect the model. Interpretability is the difference between a compliance regime based on evidence and one based on assurances. I have lived through a number of booms in which the assurances were excellent and the evidence was late.

Olah has made that argument explicitly. It moves the field from a technical curiosity to a prerequisite for the entire policy apparatus now being built around it.

A witness who cannot be coached

The argument for interpretability is essentially an argument about accountability. Every other safety mechanism operates on behaviour: filters, refusals, evaluations, monitoring. All of them can be defeated by a system that behaves well when observed. I have known a great many men of that description, and a few horses.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

Inspecting the internal computation is the only approach that does not trust the model’s outputs. That is why the work is treated as foundational rather than incremental, despite producing fewer demonstrable results than other research areas. It is the difference between questioning the witness and searching the house.

A lamp that lights one room at a time

Methods that produce interpretable features in a small model have not reliably scaled to systems with hundreds of billions of parameters and mixture-of-experts routing. Sparsity techniques have extended the reach, and the honest position is that current methods illuminate a small fraction of what a frontier model is doing. Here I will say plainly what I can and cannot vouch for: the arithmetic of the difficulty I believe, and the promised arrival date of the solution I do not.

The strategic question is whether that fraction grows fast enough to matter. If interpretability progresses linearly and capability progresses exponentially, the gap widens even as the research succeeds on its own terms. I have been wrong about this kind of race before, and I expect to be again. But I notice a thing about the arithmetic: the men building the faster engine are not the men holding the lamp, and only one of those two groups is paid to hurry.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.