Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Data center racks with glowing server lights, a sleek AI interface dashboard showing a token speed meter at 750…
Products & LaunchesBreakthroughScore: 94

OpenAI's Ultrafast Mode Hits 750 Tokens/s on GPT-5.6 Sol

OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras. Preview limited to select customers, targeting latency-sensitive enterprise workflows.

·1d ago·4 min read··16 views·AI-Generated·Report error
Share:
Source: techcrunch.comvia techcrunch_ai, hpcwireCorroborated
What is OpenAI's Ultrafast mode and how fast is GPT-5.6 Sol with it?

OpenAI launched Ultrafast, a preview mode for GPT-5.6 Sol that runs at 14x standard speed, delivering 750 output tokens per second. Powered by a Cerebras partnership, the feature targets enterprise workflows like incident response and customer service. Access is limited to a small customer group initially.

TL;DR

OpenAI previews Ultrafast mode at 14x speed · GPT-5.6 Sol outputs 750 tokens per second · Cerebras partnership powers the accelerated inference · Preview limited to small customer group initially

OpenAI launched Ultrafast mode for GPT-5.6 Sol on August 13, claiming 14x speed with 750 tokens per second. The preview, powered by Cerebras hardware, targets enterprise workloads where latency dominates cost.

Key facts

  • Ultrafast mode: 14x faster than standard GPT-5.6 Sol
  • 750 output tokens per second claimed peak throughput
  • Powered by Cerebras wafer-scale AI hardware partnership
  • Preview limited to small customer group as of August 13, 2026
  • Targets incident response, customer service, finance, e-commerce

OpenAI's new Ultrafast mode marks a shift in how the lab approaches inference speed. Rather than shrinking the model or deploying specialized variants, the company is keeping GPT-5.6 Sol's full capability while pushing throughput to 750 output tokens per second — a 14x jump over standard processing According to TechCrunch.

The feature is powered by OpenAI's partnership with chipmaker Cerebras, which specializes in wafer-scale AI accelerators. This is not a distillation or a quantization trick; it's raw hardware acceleration applied to the flagship model. The company frames the value as "more useful work per second," a phrase that signals a departure from the old tradeoff between model quality and response time.

Why latency is the new battleground

For enterprise buyers, the speed argument is straightforward: faster inference means lower cost-per-completion in high-volume workflows. OpenAI is positioning Ultrafast for incident response, customer service, financial market analysis, and e-commerce — all domains where a few hundred milliseconds of added latency translates directly into lost revenue or missed alerts.

Anthropic's Claude has a fast mode, but it doesn't deliver the same throughput. [TechCrunch notes] that Claude's accelerated option falls short of the 750 tokens per second OpenAI is claiming here. The gap matters because Anthropic has been winning enterprise credibility on coding and agentic workflows, where token throughput directly impacts task completion time.

The preview is currently limited to a small group of customers. OpenAI says it will expand access "as capacity grows," but did not disclose pricing or a general availability date. The company also did not specify which Cerebras hardware configuration underpins the mode, nor whether the 750 tokens-per-second figure is a peak or sustained rate.

The structural read

This launch is less about a new model and more about infrastructure leverage. OpenAI is effectively monetizing Cerebras's wafer-scale engineering without acquiring the chipmaker. The partnership gives OpenAI a differentiated speed claim while Cerebras gets a flagship customer proof point — a relationship that mirrors the earlier Groq-OpenAI discussions that never materialized into a deal.

What's notable is the timing. OpenAI launched GPT-5.6-Cyber just two days prior, on August 11, signaling a broader push to segment its flagship model into specialized deployment modes. Ultrafast is the performance segment; Cyber is the security segment. The strategy is to sell the same base model across different enterprise pain points rather than forcing customers to choose between speed and capability.

For developers and operators, the practical question is whether Ultrafast changes the cost equation for agentic workloads. If the 14x speedup holds in production, a Claude Code-style workflow running on GPT-5.6 Sol could complete multi-step tasks in a fraction of the wall-clock time. That would put pressure on Anthropic to either match the throughput or justify its premium on reasoning quality alone.

Key Takeaways

  • OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras.
  • Preview limited to select customers, targeting latency-sensitive enterprise workflows.

What to watch

Watch for OpenAI's next capacity expansion announcement and whether Cerebras discloses the specific hardware configuration. If Anthropic responds with a faster Claude mode or a throughput benchmark of its own, the speed war moves from marketing claims to measurable enterprise SLAs. Also track whether Ultrafast pricing lands at a premium over standard GPT-5.6 Sol tokens.

BARCELONA, CATALONIA, SPAIN - 2019/02/25: The IBM logo is seen during MWC 2019.


Source: techcrunch.com


Sources cited in this article

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

OpenAI's Ultrafast launch is a supply-chain play disguised as a feature release. By partnering with Cerebras rather than building its own custom silicon, OpenAI gets a speed advantage without the multi-year engineering cycle that Google and Amazon have committed to with TPUs and Trainium. The 750 tokens-per-second figure is impressive, but the real signal is that OpenAI is now competing on infrastructure economics, not just model intelligence. This puts Anthropic in an awkward position. Claude's fast mode is a software-level optimization, not a hardware acceleration play. If Cerebras's wafer-scale approach delivers sustained throughput at scale, Anthropic will need its own silicon partner or a fundamental improvement in inference efficiency to match the cost-per-token curve. The two-model split debate around Claude Opus 5's verbosity becomes more acute when speed is the differentiator — verbose models are expensive at high token rates. The enterprise framing is also telling. Incident response and financial market analysis are not typical ChatGPT use cases; they're mission-critical, low-latency applications where a 14x speedup justifies a premium. OpenAI is explicitly courting the same buyers that Anthropic has been winning with Claude Code and agentic workflows. The question is whether token speed alone converts those customers, or whether reasoning quality and tool-use reliability still dominate procurement decisions.
This story is part of
The AI Infrastructure War Shifts from Chips to Developer Tools
Nvidia's enterprise pivot and AWS's OpenAI bet collide with Cursor's quiet ascent
Compare side-by-side
OpenAI vs Cerebras
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all