Fei-Fei Li’s Spatial Intelligence Bet
Fei-Fei Li built the dataset that started the deep learning era. She now works on what she says comes next: machines that learn from the physical world.
I have spent a fair part of my life listening to men announce the next era, and I have learned to keep my hat on while they do it. So when Fei-Fei Li says the next era is spatial intelligence, I reach for my chair. The image dataset she assembled in the late 2000s is one of the few artefacts in modern computing that can honestly be said to have started an era: millions of photographs, labelled at a scale nobody had attempted, handed to a field that had waited years for something to point its networks at. It was not a machine. It was a heap of pictures with names on them, and it changed everything.
Her current work addresses what she sees as the limitation of everything built since. That is a bold thing to say in front of the men who built it, and she says it anyway.
The gap between words and things
Here is the trouble, and I think it is real. Language models learn from text, and text is a compressed description of the world, written by humans who already understood it before they sat down to describe it. The model learns the shorthand, not the river. Which is why a machine can explain the physics of a falling glass with perfect composure, and then watch one topple off the table with no idea what comes next.
Her argument is that spatial intelligence — understanding three-dimensional structure, distance, occlusion and physical interaction — is a separate capability that must be learned directly, and that it is the prerequisite for robotics, augmented reality and much else.
- Generating and reasoning about three-dimensional scenes rather than flat images.
- Training on sensor and video data rather than captions.
- Building representations that can be queried by a robot in real time.
The hardware problem
The obvious difficulty is that you cannot buy a dataset of physical interaction the way you can scrape a web corpus. Robots are expensive, slow and dangerous to run at scale, and simulation has a persistent gap with reality that no one has fully closed.
Li has been unusually candid that this is a longer project than the one that made her reputation. That candour is itself notable in a field where timelines are routinely compressed for fundraising purposes.
The human-facing argument
She has also been one of the more consistent voices arguing that the composition of the field matters — that a discipline whose systems will be deployed in hospitals, schools and homes should not be designed by a narrow group. She has made the point in policy settings as often as technical ones.
Why language was the easy case
Text has properties that made it unusually tractable: it is abundant, cheap to store, already digitised, and produced by humans who had pre-solved the hard representation problem. A sentence about a cup contains the concept of a cup, already abstracted. Somebody handed the machine the cup and said, here, this is a cup.
Pixels and sensor readings do not arrive with that structure. A model must discover that objects persist, that surfaces have material properties, and that a scene has depth — the things an infant learns by interacting, and there is no web crawl for physical experience. You cannot scrape a childhood.
The application case
If the capability arrives, its uses are obvious and economically significant: robots that can work in unstructured environments, augmented reality that understands a room, and simulation that reflects real physics well enough to train against.
That is also what makes the work commercially risky. Each of those applications has been five years away for a long time, and progress is gated by hardware and data collection rather than by a single algorithmic insight. Five years away is a wonderful address to live at. I have wintered there myself.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.