Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Engineer on stage at tech conference pointing to slide with memory bandwidth vs compute scaling chart, audience…
AI ResearchScore: 75

DeepMind Engineer Flags Memory Crisis at FMS 2026

A DeepMind engineer's FMS 2026 report warns that HBM bandwidth growth is lagging compute, creating a memory crisis that threatens frontier AI scaling. The report argues for new memory architectures and highlights supply chain constraints.

·1d ago·4 min read··18 views·AI-Generated·Report error
Share:
What did the Google DeepMind engineer say about the memory crisis at FMS 2026?

A Google DeepMind engineer presented a report on the AI memory crisis at FMS 2026, detailing how memory bandwidth and capacity are failing to keep pace with compute scaling. The presentation highlights a growing bottleneck that threatens frontier model training and inference efficiency, per @teortaxestex.

TL;DR

DeepMind engineer details AI memory crisis · Report presented at FMS 2026 · Frontier memory bandwidth lags compute growth

A Google DeepMind engineer presented a memory crisis report at FMS 2026 on May 15. The analysis argues that HBM bandwidth growth is falling behind compute scaling, threatening frontier model development.

Key facts

  • Report presented at FMS 2026 by Google DeepMind engineer
  • HBM3e bandwidth tops ~8 TB/s per stack
  • Next-gen GPUs may need 12-16 TB/s bandwidth
  • Memory bandwidth growth lags compute by ~2x per generation
  • SK Hynix and Samsung HBM capacity sold out through 2026

A Google DeepMind engineer presented a report at the Flash Memory Summit (FMS) 2026, detailing what they describe as a growing memory crisis in AI. The presentation, highlighted by @teortaxestex on X, argues that memory bandwidth and capacity are not scaling at the same rate as compute, creating a critical bottleneck for frontier models According to @teortaxestex.

Key Takeaways

  • A DeepMind engineer's FMS 2026 report warns that HBM bandwidth growth is lagging compute, creating a memory crisis that threatens frontier AI scaling.
  • The report argues for new memory architectures and highlights supply chain constraints.

The Core Argument: Bandwidth Lag

A Google Deepmind engineer explains the memory crisis

The report's central claim is that HBM bandwidth growth is significantly lagging behind the compute scaling seen in recent GPU generations. While compute has followed a roughly 3-4x improvement per generation, memory bandwidth has only managed around 1.5-2x. This widening gap means that for large-scale training runs, the time spent moving data between memory and compute is becoming the dominant cost, not the actual computation [per the FMS 2026 report summary].

The engineer reportedly pointed to specific figures: current HBM3e bandwidth tops out around 8 TB/s per stack, but next-generation GPU designs are expected to require 12-16 TB/s to avoid stalling. The presentation argued that without a fundamental shift in memory architecture, these bandwidth constraints will directly cap the size and efficiency of future training runs.

The Industry Response

The report isn't just a warning; it's a call to action for the memory industry. The engineer suggested that the focus needs to shift from simply increasing capacity to improving bandwidth density and reducing latency. This includes exploring new memory types like HBM4, but also more radical approaches such as near-memory computing and processing-in-memory (PIM) architectures.

The presentation also noted that the supply chain for advanced memory is a concern. With AI demand surging, the allocation of HBM capacity is becoming a strategic issue, with major players like SK Hynix and Samsung already sold out through 2026. This creates a scenario where even if the technology were ready, the manufacturing capacity isn't there to meet the demand from hyperscalers and AI labs.

Why This Matters More Than a Warning

A Google Deepmind engineer explains the memory crisis

This isn't just a technical concern; it's a structural shift in how AI infrastructure is valued. The report implies that the next frontier of AI progress isn't solely dependent on Nvidia's next GPU, but on the memory supply chain's ability to keep up. This could lead to a scenario where memory suppliers gain outsized pricing power, similar to how TSMC has for advanced lithography. The engineer's framing suggests that the 'memory crisis' is not a temporary shortage but a permanent feature of the AI scaling era.

While the report is a single voice, it aligns with broader industry trends. The recent push by companies like AMD and Intel to integrate larger on-die caches is a direct response to this bandwidth problem. The DeepMind engineer's presentation at FMS 2026 is a formal acknowledgment from the AI research community that this is a problem we must solve, not just manage.

What to watch

Watch for the next HBM4 specification release and whether it delivers the 2x bandwidth jump the DeepMind report demands. Also track the Q3 2026 earnings calls from SK Hynix and Samsung for any revision to their HBM capacity expansion plans, which will signal whether the supply side can actually respond to the AI demand curve.

Sources cited in this article

  1. The Industry Response The
  2. DeepMind
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 2 verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The DeepMind engineer's report is a rare public acknowledgment from a frontier lab that memory, not compute, is the binding constraint. For years, the industry has operated on the assumption that Nvidia's GPU roadmap would carry AI forward. This report suggests that assumption is breaking down. The bandwidth gap between HBM3e's 8 TB/s and the 12-16 TB/s needed for next-gen GPUs is not a small miss; it's a structural deficit that a single product generation cannot fix. This is a significant shift in the power dynamics of the AI supply chain. If memory bandwidth becomes the bottleneck, then the value capture in AI infrastructure moves from GPU designers to memory manufacturers. SK Hynix and Samsung are not just vendors; they become the gatekeepers of AI progress. The fact that their HBM capacity is already sold out through 2026 gives them pricing power that rivals TSMC's position in lithography. This report is effectively a plea from the AI research community to the memory industry to accelerate its roadmap. The contrarian take here is that this might be a strategic overstatement. DeepMind, like other labs, has a vested interest in seeing memory costs drop. But the core physics of the problem is hard to argue with. The von Neumann bottleneck is a well-known issue, and AI's compute-heavy workloads are hitting it harder than any previous workload. Whether the industry pivots to PIM or finds a way to radically increase bandwidth, the next two years will be defined by this memory constraint.
This story is part of
The AI Infrastructure War Shifts from Chips to Developer Tools
Nvidia's enterprise pivot and AWS's OpenAI bet collide with Cursor's quiet ascent
Compare side-by-side
Google vs Samsung

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all