A Google DeepMind engineer presented a memory crisis report at FMS 2026 on May 15. The analysis argues that HBM bandwidth growth is falling behind compute scaling, threatening frontier model development.
Key facts
- Report presented at FMS 2026 by Google DeepMind engineer
- HBM3e bandwidth tops ~8 TB/s per stack
- Next-gen GPUs may need 12-16 TB/s bandwidth
- Memory bandwidth growth lags compute by ~2x per generation
- SK Hynix and Samsung HBM capacity sold out through 2026
A Google DeepMind engineer presented a report at the Flash Memory Summit (FMS) 2026, detailing what they describe as a growing memory crisis in AI. The presentation, highlighted by @teortaxestex on X, argues that memory bandwidth and capacity are not scaling at the same rate as compute, creating a critical bottleneck for frontier models According to @teortaxestex.
Key Takeaways
- A DeepMind engineer's FMS 2026 report warns that HBM bandwidth growth is lagging compute, creating a memory crisis that threatens frontier AI scaling.
- The report argues for new memory architectures and highlights supply chain constraints.
The Core Argument: Bandwidth Lag

The report's central claim is that HBM bandwidth growth is significantly lagging behind the compute scaling seen in recent GPU generations. While compute has followed a roughly 3-4x improvement per generation, memory bandwidth has only managed around 1.5-2x. This widening gap means that for large-scale training runs, the time spent moving data between memory and compute is becoming the dominant cost, not the actual computation [per the FMS 2026 report summary].
The engineer reportedly pointed to specific figures: current HBM3e bandwidth tops out around 8 TB/s per stack, but next-generation GPU designs are expected to require 12-16 TB/s to avoid stalling. The presentation argued that without a fundamental shift in memory architecture, these bandwidth constraints will directly cap the size and efficiency of future training runs.
The Industry Response
The report isn't just a warning; it's a call to action for the memory industry. The engineer suggested that the focus needs to shift from simply increasing capacity to improving bandwidth density and reducing latency. This includes exploring new memory types like HBM4, but also more radical approaches such as near-memory computing and processing-in-memory (PIM) architectures.
The presentation also noted that the supply chain for advanced memory is a concern. With AI demand surging, the allocation of HBM capacity is becoming a strategic issue, with major players like SK Hynix and Samsung already sold out through 2026. This creates a scenario where even if the technology were ready, the manufacturing capacity isn't there to meet the demand from hyperscalers and AI labs.
Why This Matters More Than a Warning

This isn't just a technical concern; it's a structural shift in how AI infrastructure is valued. The report implies that the next frontier of AI progress isn't solely dependent on Nvidia's next GPU, but on the memory supply chain's ability to keep up. This could lead to a scenario where memory suppliers gain outsized pricing power, similar to how TSMC has for advanced lithography. The engineer's framing suggests that the 'memory crisis' is not a temporary shortage but a permanent feature of the AI scaling era.
While the report is a single voice, it aligns with broader industry trends. The recent push by companies like AMD and Intel to integrate larger on-die caches is a direct response to this bandwidth problem. The DeepMind engineer's presentation at FMS 2026 is a formal acknowledgment from the AI research community that this is a problem we must solve, not just manage.
What to watch
Watch for the next HBM4 specification release and whether it delivers the 2x bandwidth jump the DeepMind report demands. Also track the Q3 2026 earnings calls from SK Hynix and Samsung for any revision to their HBM capacity expansion plans, which will signal whether the supply side can actually respond to the AI demand curve.









