The Memory Wall: Why High-Bandwidth Memory Became the Bottleneck
An accelerator is limited less by how fast it computes than by how fast it can be fed. The stack of memory bolted to it became the tightest constraint in the supply chain.
The performance of a machine learning accelerator is limited less by how fast it can compute than by how fast it can be fed. Every arithmetic operation needs data, and if the memory cannot supply that data quickly enough, the processor sits idle, burning power and accounting for it as utilisation.
Solving that problem requires memory placed physically close to the processor and connected by very wide, very short pathways. That is what high-bandwidth memory provides, and it became the tightest constraint in the entire supply chain.
Taller stacks, worse yields
- The memory dies are stacked vertically and joined with microscopic connections.
- Yields fall sharply as the stack grows taller, and customers demand near-perfect parts.
- Only a small number of manufacturers have mastered the process.
- The manufacturing capacity is shared with conventional memory, which competes for the same lines.
You cannot retrofit a short stack
Because memory capacity per accelerator is fixed at design time, a shortage cannot be solved by adding more memory to an existing product. It constrains the number of complete units that can ship, regardless of how many processor dies are available.
That made memory allocation a board-level concern at every major buyer, and it explains why so much of the recent capital spending announcements mention packaging and memory capacity alongside logic.
Bandwidth improves, appetite improves faster
Memory bandwidth has improved dramatically, but so has the appetite of models for it, and the appetite has been winning. Whether the industry closes the gap depends on packaging advances, on new memory types, and on algorithmic work that reduces the amount of data a model must move. All three are active areas, and none of them is guaranteed.
One bad layer ruins the assembly
Stacking memory dies requires connecting thousands of microscopic pads between layers, then attaching the finished stack to the processor through a silicon interposer. A single defective layer ruins the whole assembly, and the assemblies are large enough that yield losses are expensive.
The number of manufacturers who can produce these stacks at volume is small, and they must simultaneously supply conventional memory for phones, computers and cars. Capacity decisions made for one market constrain the other.
Algorithms redesigned around a parts shortage
Because memory is the constraint, model architectures are increasingly designed around it rather than around raw compute. Mixture-of-experts routing, which activates only part of a model per token, exists partly to reduce the amount of data that must be moved per unit of work.
An algorithm shaped by what a parts supplier can ship, rather than by a performance target somebody chose, is an unusual direction for a field that usually describes itself as idea-led.
Image credit and licence details for every photograph on this site are listed on the credits page. This article is editorial content; it carries no sponsored material.