Google's 7- and 8-year-old TPUs still run at 100% utilization, according to Amin Vahdat, VP and GM of AI and Infrastructure at Google Cloud. He cites Jevons Paradox: efficiency gains in AI hardware drive explosive demand, not reduced usage.
Key facts
- Google's 7-8 year old TPUs at 100% utilization.
- Amin Vahdat: VP/GM of AI and Infrastructure, Google Cloud.
- Jevons Paradox cited as demand driver.
- Source: @rohanpaul_ai via a16z YouTube video.
- Google's 2025 capex: $75B.
Amin Vahdat, Google Cloud's VP and GM of AI and Infrastructure, said on an a16z YouTube video that Google's 7- and 8-year-old TPUs are still seeing 100% utilization According to @rohanpaul_ai. The claim, relayed by @rohanpaul_ai, underscores a counterintuitive trend in AI infrastructure: older chips remain fully saturated despite newer generations being deployed.
Vahdat attributes this to Jevons Paradox, the economic principle that as a technology becomes more efficient, its consumption increases rather than decreases. "Once something gets more efficient, its use just explodes," he said. "AI is sitting right there now." This explains why Google continues to operate legacy TPUs at full capacity, even as it rolls out newer TPU generations like v6 and v7.
The observation challenges the common assumption that hardware upgrades retire older silicon. Instead, efficiency gains in AI models and training algorithms are expanding the total compute demand, keeping every available chip busy. This dynamic has direct implications for capacity planning: cloud providers like Google, AWS, and Microsoft are racing to add data centers, but existing fleets—even aging ones—remain critical assets.
Google does not publicly disclose the exact utilization rates of its TPU fleet, but Vahdat's statement aligns with industry reports of AI compute shortages persisting through 2025 and into 2026 [Reuters has reported on sustained GPU and TPU demand]. The company's own infrastructure investments, including $75B in capital expenditures for 2025, reflect this unrelenting demand.
The practical takeaway: don't count out old hardware. For enterprises planning AI workloads, the lesson is that efficiency gains in models—like those from quantization or distillation—will not free up capacity; they will attract more users and workloads, keeping utilization high across the entire fleet.
What to watch
Watch for Google Cloud's next earnings call or infrastructure blog to see if Vahdat quantifies TPU fleet utilization or discloses how many legacy TPU generations remain in production. Also track whether hyperscalers announce extended lifespans for older accelerators amid persistent AI compute shortages.






