Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

AI Agents: The Long Road From Demo to Deployment

The gap between an agent that works once and an agent that works reliably is where most projects die. The arithmetic is unforgiving and the demos never show it.

An automated workflow on a screen — photo by Tsinkala, licensed under CC BY-SA 4.0 via Wikimedia Commons.

The most compelling demonstrations in artificial intelligence are agentic: a system is given a goal, it breaks the goal into steps, uses tools, recovers from errors and finishes the task. In a recorded demo, this is remarkable.

In production, the same system produces a different set of numbers. Reliability compounds. A process with twenty steps, each of which succeeds 95 percent of the time, completes correctly roughly a third of the time.

Errors cluster in five places

  • Error recovery: the agent escalates or loops instead of trying an alternative.
  • Ground truth: it acts confidently on a wrong assumption about the environment.
  • Permissions: it has more access than it needs, and occasionally uses it.
  • Cost: a long agent run can consume an unpredictable amount of inference.
  • Observability: nobody can explain why a particular run went wrong.

One bounded workflow, a human on the consequential steps

The deployments that succeed are narrower than the demos. They handle one bounded workflow, with a human approving consequential actions, a hard limit on the number of steps, and detailed logging. The value comes from removing the tedious middle of a process, not from automating an entire job.

Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.

The most useful engineering insight from the last few years is that an agent is a distributed system, and it should be built with the caution that implies: idempotent operations, verification at the boundaries, timeouts, and rollback paths.

Open-ended autonomy is still a research problem

Full autonomy on open-ended tasks remains an unsolved research problem. Narrow autonomy on well-specified tasks is available now and improving quickly. The commercial opportunity in the near term is overwhelmingly in the second category, and the companies that understand the difference are the ones shipping.

A test that passes, a schema that validates

A reliable agent needs a way to check its own intermediate results. That means the environment must expose verifiable signals — a test that passes, a schema that validates, a database that confirms — and designing those signals is engineering work that has nothing to do with the model.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

Systems without verification rely on the model’s judgement about whether it succeeded, which is the weakest available signal and the one most likely to be confidently wrong. The best-performing deployments invest heavily in making failure visible rather than in making the agent smarter.

Fifty steps can cost a hundred responses

An agent that takes fifty steps may consume a hundred times the inference of a single response. That is affordable for a task that replaces an hour of human work and unaffordable for a task that replaces two minutes.

The implication is that the near-term market for agents is high-value, low-volume work: financial analysis, code migration, compliance review, research synthesis. Consumer-facing agents face a much harder economic case, which is why so few of them have shipped.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.