TileRT AI's TILERT software claims 1.9x decode interactivity on NVIDIA Blackwell GPUs at the same per-token cost, per @SemiAnalysis_. The gain, if real, pressures Groq, Cerebras, and SambaNova's purpose-built inference chips.
Key facts
- 1.9x decode interactivity boost claimed on NVIDIA Blackwell
- Same per-token cost, no hardware change
- Software-only optimization from TileRT AI
- Targets Groq, Cerebras, SambaNova latency advantage
- No benchmark methodology disclosed
TileRT AI's TILERT claims to boost decode interactivity by 1.9x on the same NVIDIA Blackwell GPUs at the same per-token cost, according to a tweet from @SemiAnalysis_ per the tweet. The claim is software-only, meaning no hardware changes, and it directly targets the latency advantage that purpose-built inference chips like Groq, Cerebras, and SambaNova have marketed. If TILERT's 1.9x claim holds in production, it could erode the economic case for buying specialized inference hardware for many workloads.
Key Takeaways
- TileRT AI's TILERT claims 1.9x decode interactivity on NVIDIA Blackwell at same cost, pressuring Groq/Cerebras/SambaNova.
- Claim unverified, but if real, erodes custom-chip value.
Why this matters for the inference market

The entire value proposition of Groq, Cerebras, and SambaNova rests on decode speed — the per-token generation latency that determines how "interactive" an LLM feels. These vendors sell custom silicon and memory architectures to beat NVIDIA's GPUs on this metric. A pure-software optimization that delivers 1.9x on existing Blackwell hardware, at no extra cost, undercuts that pitch. The tweet frames it as an alert, suggesting the performance gap may be closing faster than the specialized vendors have projected.
What TILERT actually is
TileRT AI, the company behind TILERT, has not published technical details. The tweet does not specify the optimization technique — whether it's a better KV-cache management, a new batching schedule, or a kernel-level tweak. The 1.9x figure is presented without benchmark methodology or reproducibility details. TileRT did not disclose the specific techniques or benchmarks used to achieve the 1.9x figure, leaving the claim unverified outside the vendor's own testing. This is typical of pre-print announcements, but it means the number should be treated as a claim, not a result.
The strategic read
For NVIDIA, TILERT is a defensive win: it keeps workloads on Blackwell rather than ceding them to custom chips. For Groq, Cerebras, and SambaNova, the threat is existential if the claim generalizes. They have argued that GPUs hit a latency wall; TILERT suggests that wall is software-addressable. The real test will be third-party reproduction on standard Blackwell deployments, not vendor-published numbers. Watch for independent benchmarks from cloud providers or MLPerf submissions that include TILERT-optimized results.
What to watch
Watch for third-party reproduction of TILERT's 1.9x claim on standard Blackwell deployments, ideally via MLPerf or a cloud provider's public benchmark. If Groq or Cerebras respond with updated latency numbers in the next quarter, that signals real competitive pressure. Also track TileRT's technical disclosure — a paper or kernel release would substantiate the claim.








