OpenAI launched Ultrafast mode for GPT-5.6 Sol on August 13, claiming 14x speed with 750 tokens per second. The preview, powered by Cerebras hardware, targets enterprise workloads where latency dominates cost.
Key facts
- Ultrafast mode: 14x faster than standard GPT-5.6 Sol
- 750 output tokens per second claimed peak throughput
- Powered by Cerebras wafer-scale AI hardware partnership
- Preview limited to small customer group as of August 13, 2026
- Targets incident response, customer service, finance, e-commerce
OpenAI's new Ultrafast mode marks a shift in how the lab approaches inference speed. Rather than shrinking the model or deploying specialized variants, the company is keeping GPT-5.6 Sol's full capability while pushing throughput to 750 output tokens per second — a 14x jump over standard processing According to TechCrunch.
The feature is powered by OpenAI's partnership with chipmaker Cerebras, which specializes in wafer-scale AI accelerators. This is not a distillation or a quantization trick; it's raw hardware acceleration applied to the flagship model. The company frames the value as "more useful work per second," a phrase that signals a departure from the old tradeoff between model quality and response time.
Why latency is the new battleground
For enterprise buyers, the speed argument is straightforward: faster inference means lower cost-per-completion in high-volume workflows. OpenAI is positioning Ultrafast for incident response, customer service, financial market analysis, and e-commerce — all domains where a few hundred milliseconds of added latency translates directly into lost revenue or missed alerts.
Anthropic's Claude has a fast mode, but it doesn't deliver the same throughput. [TechCrunch notes] that Claude's accelerated option falls short of the 750 tokens per second OpenAI is claiming here. The gap matters because Anthropic has been winning enterprise credibility on coding and agentic workflows, where token throughput directly impacts task completion time.
The preview is currently limited to a small group of customers. OpenAI says it will expand access "as capacity grows," but did not disclose pricing or a general availability date. The company also did not specify which Cerebras hardware configuration underpins the mode, nor whether the 750 tokens-per-second figure is a peak or sustained rate.
The structural read
This launch is less about a new model and more about infrastructure leverage. OpenAI is effectively monetizing Cerebras's wafer-scale engineering without acquiring the chipmaker. The partnership gives OpenAI a differentiated speed claim while Cerebras gets a flagship customer proof point — a relationship that mirrors the earlier Groq-OpenAI discussions that never materialized into a deal.
What's notable is the timing. OpenAI launched GPT-5.6-Cyber just two days prior, on August 11, signaling a broader push to segment its flagship model into specialized deployment modes. Ultrafast is the performance segment; Cyber is the security segment. The strategy is to sell the same base model across different enterprise pain points rather than forcing customers to choose between speed and capability.
For developers and operators, the practical question is whether Ultrafast changes the cost equation for agentic workloads. If the 14x speedup holds in production, a Claude Code-style workflow running on GPT-5.6 Sol could complete multi-step tasks in a fraction of the wall-clock time. That would put pressure on Anthropic to either match the throughput or justify its premium on reasoning quality alone.
Key Takeaways
- OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras.
- Preview limited to select customers, targeting latency-sensitive enterprise workflows.
What to watch
Watch for OpenAI's next capacity expansion announcement and whether Cerebras discloses the specific hardware configuration. If Anthropic responds with a faster Claude mode or a throughput benchmark of its own, the speed war moves from marketing claims to measurable enterprise SLAs. Also track whether Ultrafast pricing lands at a premium over standard GPT-5.6 Sol tokens.

Source: techcrunch.com









