Alexandr Wang and Scale AI: The Data Army
The least glamorous layer of the AI stack turned out to be one of the most defensible businesses in it, and Alexandr Wang got there first.
Every conversation about artificial intelligence eventually runs into a question that is harder than it sounds: where did the training data come from, and who checked it? The men who ask this question are usually the ones not being paid to answer it.
Alexandr Wang built one of the largest companies in the field by answering that question at industrial scale. His firm supplies labelled and curated data — text, images, audio, sensor readings and, increasingly, the expert demonstrations used to fine-tune models. The least romantic sentence in this column.
An underrated moat
The assumption in the early years was that data labelling was a commodity: cheap, low-skill and easily replicated. That assumption did not survive contact with frontier research. Modern model development needs domain specialists — physicians, lawyers, mathematicians, speakers of hundreds of languages — working to consistent standards.
Assembling and managing that workforce is an operations problem, not a research problem, and it is exactly the kind of problem that technology companies are bad at. A firm that can hire brilliant people and cannot schedule them has an expensive waiting room.
- Quality control across a distributed workforce numbering in the hundreds of thousands.
- Security and confidentiality requirements from defence and healthcare clients.
- Sourcing experts fast enough to keep up with model release schedules.
The strategic position
Sitting between the labs and the raw material of their models is a powerful place to be — the position of the ferryman at a river with one crossing. Every frontier lab is a customer. Every new capability — video, robotics, agentic workflows — creates a new data requirement, and a new contract.
The risk is disintermediation. Labs have invested heavily in synthetic data and automated evaluation, both of which reduce dependence on human labelling — and a great many middlemen have been retired by a cleverer machine, though rarely as quickly as announced. Wang’s response has been to move up the value chain into evaluation, model development services, and government work.
The frontier of specialist data
The fastest-growing part of the market is not generic labelling but expert work: demonstrations from surgeons, corrections from mathematicians, evaluations written by specialists in narrow fields. This data cannot be produced by a large low-cost workforce, because the supply of qualified people is the constraint — you cannot hire a thousand surgeons by Monday.
That shift changes the economics of the business. Specialist annotation carries much higher margins and much higher barriers, and it is harder for a technology company to replicate with automation — the sort of sentence I would not have written about a labelling shop ten years ago.
The government market
Defence and public sector contracts are a substantial and growing part of the business. They come with security requirements, long procurement cycles and political exposure that commercial work does not, and they are correspondingly harder to displace. A customer who takes two years to buy takes longer to leave.
They also create reputational questions the company has had to navigate publicly. Being a supplier to both leading labs and national defence establishments puts a firm where it will eventually be asked whose interests it serves. A man who supplies both ends of a quarrel has bought himself a long autumn of explaining.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.