Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

Why AI Watermarks Keep Failing the People Who Need Them

Almost everyone agrees that people should be able to tell whether a piece of media was generated by a machine. Almost nobody agrees on how to make that happen.

Digital image processing on a screen — photo by Ngwedi Mokgoro, licensed under CC BY-SA 4.0 via Wikimedia Commons.

Almost everyone agrees that people should be able to tell whether a piece of media was generated by a machine. Almost nobody agrees on how to make that happen.

Three techniques, and what breaks each of them

The first approach embeds a signal in the content itself — an invisible pattern in the pixels or a statistical bias in generated text. These marks can be robust against compression and cropping, but they are trivially removed by anyone willing to run the output through a re-encoder, and open-weight models can be fine-tuned to omit them.

The second approach attaches metadata at the point of creation, following a standard that records origin and edit history. This survives formats that respect it and disappears at the first screenshot.

The third approach is a registry: the generating service records what it produced and publishes a lookup. It works only if the platform is honest, which is not guaranteed.

Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.
  • Invisible marks: removable, and false positives are a real problem.
  • Metadata: loses fidelity as soon as the file is re-shared.
  • Registries: depend on voluntary participation by the producer.

The signal is weakest exactly where it is most needed

The deeper issue is that provenance signals are most needed on the platforms where they are least enforced. A newsroom may carefully preserve metadata; a messaging app will strip it by default; a screenshot removes everything.

Layered signals, treated as evidence rather than proof

The most credible deployments combine weak signals, and rely on them as one input among several rather than as proof. Detection models, editorial verification, source reputation and platform policy all contribute. That is unsatisfying to anyone hoping for a technical fix, but it matches how authentication has worked in every other medium.

When the detector accuses a human

Invisible watermarking has a failure mode that receives less attention than removal: content produced by humans can be flagged as machine-generated. A statistical detector operating at scale will produce false accusations, and an accusation of fabrication is damaging even when it is wrong.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

That risk is why detection evidence is rarely treated as conclusive. Used as one signal among several, it is useful. Used as proof, it is a liability.

Record the origin instead of hunting the output

The more robust approach inverts the default: instead of trying to detect generated content after the fact, record the origin of everything and treat unverifiable material as unverified. This is how the standard for content credentials works, and how photography authentication has long operated.

Its weakness is adoption. A signing scheme is useful only if the cameras, editors and publishing tools that matter all participate, and platforms must then actually display the result. Both conditions are currently only partly met.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.