The Pause That Wasn’t: How Anthropic Quietly Removed Its Own Brake
The commitment that made Anthropic’s warnings credible was the first thing to go when it became expensive. The company’s reasons are on the record, and they are not all bad.
When Anthropic published the first version of its Responsible Scaling Policy in September 2023, it was the first frontier lab to write down, in public, the conditions under which it would stop. The scheme was simple enough to explain in a paragraph. Define capability thresholds. Define the safeguards each threshold requires. Commit not to train a system unless the safeguards are in place first. If capabilities outrun safeguards, pause.
Eleven other companies adopted variants of it. It became the working template for what a frontier AI safety framework looks like, and it was cited in the construction of California’s SB-53, New York’s RAISE Act and the European Union’s AI Act. For two and a half years it was the single most consequential voluntary commitment in the industry, and the reason was precisely that it had teeth: it bound the company in advance, before it knew what it would want to do.
What version 3.0 changed
On 24 February 2026 Anthropic published RSP v3.0. The categorical pause is gone. In its place is a dual condition: development is delayed only if Anthropic is judged to be leading the AI race and catastrophic risk is judged material. Both must hold simultaneously. Some of the old unilateral commitments were reclassified as “industry-wide recommendations,” which the company says it will “strive to advance” without committing to them unconditionally.
Two weeks before that, Mrinank Sharma, who led Anthropic’s safeguards research, resigned. His letter did not name the policy change, but it named the condition: “Throughout my time here, I’ve repeatedly seen how hard it is to truly let our values govern our actions… we constantly face pressures to set aside what matters most.”
Anthropic’s reasons, in its own words
The company gave two, and both are argued rather than asserted.
The collective action problem. Catastrophic risk depends on what all frontier developers do, not on what one of them does. If Anthropic slows to implement mitigations while competitors do not, the competitors take the lead — and the result, on Anthropic’s account, is worse for safety, because the company loses the capacity to do safety research and to advance the public benefit at all.
Perverse incentives. Under the old policy, crossing a declared capability threshold triggered very costly mitigations or a pause. That gave the company a reason not to declare thresholds it had in fact crossed, which could produce a misleading picture of the risk. A rule that penalises honesty about capability is a bad rule.
Co-founder and chief scientific officer Jared Kaplan put the retrospective case to TIME more bluntly: the belief that a clean line could be drawn between “dangerous” and “safe” had been a “naive” idea, and with competitors going full speed ahead, “our making strict unilateral commitments isn’t realistic.”
What was added
It is not accurate to describe v3.0 as pure subtraction. Two new instruments came in.
- Risk Reports, published every three to six months, which go beyond the existing system cards by giving an overall risk assessment rather than a per-model one.
- A Frontier Safety Roadmap: detailed, public, explicitly non-binding safety goals across security, alignment, safeguards and policy, against which Anthropic says it will openly grade its own progress.
- A commitment to match a competitor’s mitigations where they are more effective and can be implemented at comparable cost.
The Governance Institute’s published assessment moved from negative to cautiously positive on closer reading, while keeping the central objection: dropping the pause makes it more likely that Anthropic deploys a model posing unacceptable risk. Its conclusion is worth quoting, because it is the fairest sentence anyone has written about this change: it is “better to be honest about constraints than to keep commitments that won’t be followed in practice.”
Why it still matters
Honesty about constraints is a virtue. But there is a difference between admitting you would probably have broken a promise and declining to make it. The old pause was valuable precisely because it was uncomfortable — because it bound the company at a moment when no one could know which way the competitive pressure would run. A dual condition that only binds when you are already winning and the risk is already obvious binds on the two occasions when restraint is easiest.
That is the whole of the critique. It does not require bad faith. It requires only that a company, facing a $200 million Pentagon contract and rivals who were not slowing down, re-priced a commitment — and that the commitment it chose to re-price was the one the public had been told was load-bearing.
See the timeline for how this sits against the warnings of 2023, and the receipts for the sources.
Sources for every claim in this article are dated and listed on the receipts page. Image credit and licence details are on the credits page. This article is editorial content; it carries no sponsored material.