SemiAnalysis and Supermicro's Vik Malyala declared storage the new AI bottleneck. Between KV cache offload demand and surging SSD costs, storage has shifted from afterthought to headliner in AI infrastructure discussions.
Key facts
- KV cache offload drives storage demand in AI inference.
- Supermicro's Petascale solutions balance density and IO bandwidth.
- SSD prices are rising, making storage cost the new concern.
- Partners include VAST, WEKA, DDN, Qumulo, OSNEXUS.
- E1.S and E3.S form factors enable higher density in 1U/2U.
The AI infrastructure conversation has pivoted from compute to storage. According to @SemiAnalysis_, Vik Malyala of Supermicro walked through the company's evolving storage stack, emphasizing that "storage, whether it be for KV cache offload or just the incredible rise in cost of SSDs, is the primary thing to focus on for a lot of users."
The Form Factor Shift
Malyala traced the evolution: "Initially it started with the form factors. You have U.2s that were popular, or still are popular for that matter. Then you have E1.S and E3.S drives, mainly because of the density that we can bring in a 1U or 2U form factor." Supermicro's "Petascale" solutions aim to balance storage density with IO bandwidth, a critical trade-off as inference workloads generate massive KV caches that must be offloaded from HBM to SSD to keep costs under control.
Software-Defined Storage Partnerships

"Hardware is only one part of it," Malyala noted. Supermicro is working with software-defined storage vendors including VAST, WEKA, DDN, Qumulo, and OSNEXUS. This allows customers to deploy flexible storage architectures rather than being locked into traditional SAN/NAS approaches. The implication: as GPU clusters scale to tens of thousands of accelerators, the storage layer must keep pace not just in capacity but in bandwidth, especially for checkpointing and inference cache management.
The unique take here is that KV cache offload is now a first-order design constraint for AI infrastructure, not a niche optimization. As context windows grow (Claude 3.5 Sonnet supports 200K tokens; Gemini 1.5 Pro supports 2M), the memory required to store attention keys and values per request scales linearly with sequence length and batch size. Offloading that cache to SSDs reduces HBM pressure but requires storage with low latency and high throughput — precisely the gap Supermicro's Petascale solutions target.
What to watch
Watch for Supermicro's next quarterly earnings call (expected late July 2026) for Petascale storage revenue breakdown, and for benchmark comparisons from MLPerf Inference v4.0 that measure KV cache offload latency impact on end-to-end performance.








