Global Edition
The Doom Ledger
Est. 2026
AI is not a new church, and people don’t need a new pope.

Grok, ChatGPT, Claude, Gemini: A Field Guide to the Frontier

Marketing pages all claim the same thing. The differences that matter to real users show up elsewhere.

A chat interface on a screen — photo by Pneuma, commissioned by the Wikimedia Foundation, licensed under CC BY-SA 4.0 via Wikimedia Commons.

Every major assistant claims to be the most capable, the safest and the most useful. The claims are not exactly false; they are unfalsifiable, because each company chooses the evaluation on which it performs best.

For anyone actually choosing a tool, the useful distinctions are in places the marketing pages do not cover.

Six things to compare instead

  • Refusal behaviour: how often the system declines a legitimate request, and how gracefully.
  • Context handling: what happens with a long document or a long conversation.
  • Tool use: whether it can search, run code and call services reliably.
  • Latency: the difference between an answer in two seconds and one in twenty seconds.
  • Data terms: what happens to your input, and whether you can turn it off.
  • Ecosystem: whether it works where your work already lives.

Standardised tasks in ideal conditions

Public benchmarks measure performance on standardised tasks in ideal conditions. They say little about whether a model will format output the way your system needs, how it behaves when a tool call fails, or whether it hallucinates citations in an obscure domain. Every organisation that has deployed these systems has developed its own private evaluation, and that evaluation is usually the only one it trusts.

Advertisementin-article · responsiveAfter the opening section of a long article. Never between a heading and its own body.

Twenty tasks of your own are better than any leaderboard

The productive method is to build a small set of your own tasks — twenty is enough — with known correct answers, and run every candidate against it. Include the awkward cases: ambiguous instructions, missing information, requests the system should refuse. The results will not match any published leaderboard, and they will be far more useful.

Ask about failed tool calls and get no number

Ask a provider how often its model refuses a legitimate request and you will not get a number. Ask how it behaves when a tool call fails, or what proportion of citations are fabricated in a niche domain, and the answers are similarly unavailable.

Advertisementin-article-2 · responsiveRoughly two thirds down a long article.

Those are the metrics that determine whether a deployment succeeds, and none of them appear on a model card. The reason is not conspiracy but measurement cost: producing them requires an evaluation suite tailored to each customer’s domain, which no vendor can build in advance.

A year of tuned prompts is the real lock-in

In practice, most organisations use more than one provider. The cost of switching is low at the start and rises as prompts, retrieval pipelines and evaluation suites accumulate around a particular model’s idiosyncrasies.

That accumulation is the real lock-in, more than pricing or contractual terms. A team that has spent a year tuning prompts for one system will not migrate for a marginal capability improvement, which is why providers invest so heavily in tooling and integration.

Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.

Related

AI Agents: The Long Road From Demo to Deployment

Reasoning Models, Explained Without the Marketing

Advertisementfooter-banner · 970x90End of page, above the site footer. Never inside the footer itself.