Grok, ChatGPT, Claude, Gemini: A Field Guide to the Frontier
Marketing pages all claim the same thing. The differences that matter to real users show up elsewhere.
Every major assistant claims to be the most capable, the safest and the most useful. The claims are not exactly false; they are unfalsifiable, because each company chooses the evaluation on which it performs best.
For anyone actually choosing a tool, the useful distinctions are in places the marketing pages do not cover.
Six things to compare instead
- Refusal behaviour: how often the system declines a legitimate request, and how gracefully.
- Context handling: what happens with a long document or a long conversation.
- Tool use: whether it can search, run code and call services reliably.
- Latency: the difference between an answer in two seconds and one in twenty seconds.
- Data terms: what happens to your input, and whether you can turn it off.
- Ecosystem: whether it works where your work already lives.
Standardised tasks in ideal conditions
Public benchmarks measure performance on standardised tasks in ideal conditions. They say little about whether a model will format output the way your system needs, how it behaves when a tool call fails, or whether it hallucinates citations in an obscure domain. Every organisation that has deployed these systems has developed its own private evaluation, and that evaluation is usually the only one it trusts.
Twenty tasks of your own are better than any leaderboard
The productive method is to build a small set of your own tasks — twenty is enough — with known correct answers, and run every candidate against it. Include the awkward cases: ambiguous instructions, missing information, requests the system should refuse. The results will not match any published leaderboard, and they will be far more useful.
Ask about failed tool calls and get no number
Ask a provider how often its model refuses a legitimate request and you will not get a number. Ask how it behaves when a tool call fails, or what proportion of citations are fabricated in a niche domain, and the answers are similarly unavailable.
Those are the metrics that determine whether a deployment succeeds, and none of them appear on a model card. The reason is not conspiracy but measurement cost: producing them requires an evaluation suite tailored to each customer’s domain, which no vendor can build in advance.
A year of tuned prompts is the real lock-in
In practice, most organisations use more than one provider. The cost of switching is low at the start and rises as prompts, retrieval pipelines and evaluation suites accumulate around a particular model’s idiosyncrasies.
That accumulation is the real lock-in, more than pricing or contractual terms. A team that has spent a year tuning prompts for one system will not migrate for a marginal capability improvement, which is why providers invest so heavily in tooling and integration.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.