The Safety Premium: What Caution Is Actually Worth
Safety positioning was Anthropic’s moat. The question after February 2026 is whether the moat survived the decision that made it.
There is a moment in every enterprise procurement process where the technical evaluation stops being decisive and the risk question takes over. It usually happens in a meeting with people who will not read the model card. They want to know one thing: if this goes wrong, can we show that we chose carefully?
That question has an answer that has very little to do with safety and a great deal to do with documentation. And for most of the last three years, Anthropic had the best documentation in the market.
How the premium is built
Three mechanisms, none of them sinister.
- Justifiability. A buyer who must defend a vendor choice to a board, a regulator or an auditor prefers the vendor whose process can be described in a paragraph. The Responsible Scaling Policy gave procurement teams that paragraph.
- Transferability. When eleven other companies adopted comparable frameworks and the framework fed into SB-53, the RAISE Act and the EU AI Act, buying from the company that wrote the template became the low-risk default.
- Seriousness. In a category where every competitor promises transformation, the vendor that volunteers its own limits reads as the adult in the room. That impression is worth more than any benchmark.
What it is worth, and what it costs
The premium shows up in enterprise share, in regulated industries, and in the ability to charge without being compared line-by-line on capability. It also has a cost structure that is easy to miss: a safety brand is an asset that has to be continuously serviced. Every published commitment is a hostage to the next commercial decision.
Which is what makes February 2026 the important date in this file. When Anthropic removed the categorical pause commitment from RSP v3.0, it did not merely change a policy. It reduced the value of the thing enterprise buyers were paying for. The dual condition that replaced it binds only when the company is already leading and the risk is already obvious — which is to say, it binds when restraint costs least.
The safeguards researcher who resigned two weeks earlier described the pressure directly. The co-founder who later called the old certainty “naive” was not being cynical; he was being accurate. Both statements are about the same fact: the commitment was expensive, and the expense had arrived.
What would restore it
Two things, and the company has begun both.
Permanent access for independent evaluators, announced in September 2026, is the real article. A safety brand that cannot be externally checked is a marketing position; one that can be is a governance arrangement. The difference is whether the auditor reports to the company or to the public.
The second is duller and more important: publish the Risk Reports on schedule, grade yourself against the Frontier Safety Roadmap with the same candour you apply to describing the risks, and accept that a bad grade is the price of the premium. If the grading turns out to be generous, the premium goes with it.
The wider point
Every company in this industry now sells a version of reassurance. The useful question is never whether the reassurance is sincere — sincerity is cheap and abundant. It is whether the reassurance is checkable, and by whom.
That is the standard this publication will apply, to Anthropic and to anyone else who asks to be trusted with the definition of the risk.
Sources for every claim in this article are dated and listed on the receipts page. Image credit and licence details are on the credits page. This article is editorial content; it carries no sponsored material.