Yann LeCun Still Thinks Large Language Models Are a Dead End
The field’s most persistent dissenter has been making the same argument for years. He may yet be right.
Yann LeCun won a Turing Award for work that made modern deep learning possible. He then spent the years in which that work produced its most commercially successful results arguing that the results are a detour. I have a considerable respect for a man who will stand up at his own banquet and tell the guests the soup is wrong.
His position is that predicting the next token in a sequence is not the same thing as understanding the world — and that no amount of scale will close the difference. Systems trained this way, he argues, are extraordinarily good at manipulating language and correspondingly poor at reasoning about physical reality.
The alternative he proposes
LeCun’s proposed direction is a family of architectures built around world models: systems that learn a compressed internal representation of how their environment behaves, and then plan within that representation. The appeal is that such systems could learn from observation rather than from exhaustively labelled data, closer to how an infant learns that objects persist when hidden. I have never seen an infant demand a labelled dataset, and I have seen many learn the rules of a room before they could name a single object in it.
- Learning from video and sensor streams rather than text alone.
- Planning in a learned latent space instead of predicting tokens.
- Architectures with explicit memory and hierarchical control.
Why the argument is uncomfortable
The uncomfortable part of LeCun’s critique is that it implies the tens of billions of dollars being spent on scaling current architectures may not compound the way investors expect. That is not a popular position among people deploying that capital. There is a particular silence that falls over a room when somebody mentions the size of the stake.
It is possible for both sides to be right in their own terms: current systems may be extremely valuable products while still being a poor path to the kind of general capability their builders claim to be pursuing. A horse that wins the county fair is a fine horse, which is not the same as the claim printed on the programme.
The case for taking him seriously
LeCun has been on the unfashionable side of this argument before. When he and a small group of colleagues insisted that learning representations from data would beat hand-engineered features, almost nobody agreed. The field changed its mind. Whether it changes its mind again is the most interesting open bet in the discipline.
What would settle the argument
The disagreement is empirical, which means it is resolvable in principle. If scaling current architectures continues to produce qualitative new capabilities, the sceptical position weakens. If progress flattens as compute grows, it strengthens.
The difficulty is that both sides can explain any observation after the fact. A plateau becomes a data problem; a breakthrough becomes evidence that the architecture was sufficient all along. The field needs predictions that can fail, and it does not have many.
The institutional dimension
There is a practical reason the argument matters beyond philosophy. Capital allocation follows belief. If world models are the right path, then money spent on ever-larger language models is partly misallocated, and the companies with the strongest position today may be building the wrong expertise.
LeCun has continued to work on his alternative in an industry lab, publishing methods and releasing code. That is a more substantive contribution than the critique itself, and it is the part of his position that will ultimately be judged.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.