Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

A glossy dark AI accelerator chip labeled Jalapeño sits on a test bench beside a green pepper, with benchmark charts…

OpenAI Jalapeño Chip Beats Nvidia Blackwell on InferenceX

OpenAI's Jalapeño chip beat Nvidia Blackwell on InferenceX at Hot Chips 2026, with more tokens per watt. Volume deployment slips to 2027, raising questions about the competitive window.

·11h ago·4 min read··40 views·AI-Generated·Report error
Share:
Source: techcrunch.comvia techcrunch_ai, openai_blog, next_big_future, the_decoder, @SemiAnalysis_, @intheworldofai, @tomshardwareMulti-Source
How does OpenAI's Jalapeño chip perform against Nvidia's Blackwell on the InferenceX benchmark?

OpenAI's Jalapeño inference chip beat Nvidia's Blackwell system on SemiAnalysis' InferenceX benchmark at Hot Chips 2026, delivering more tokens per user and more throughput per kilowatt. Richard Ho, OpenAI's head of hardware, called the results a 'very, very significant performance advance' over state of the art. Small deployment begins late 2026, with scale in 2027.

TL;DR

Jalapeño beats Blackwell on SemiAnalysis InferenceX benchmark · Full deployment slips to 2027, small volumes late 2026 · Chip co-designed with Broadcom for prefill, KV cache locality

OpenAI's Jalapeño inference chip beat Nvidia's Blackwell on SemiAnalysis' InferenceX benchmark at Hot Chips 2026. Richard Ho, OpenAI's head of hardware, called the results a "very, very significant performance advance over state of the art."

Key facts

  • Jalapeño beats Nvidia Blackwell on SemiAnalysis InferenceX benchmark
  • More tokens per user and throughput per kilowatt than state-of-the-art
  • Small deployment late 2026, significant volume in 2027
  • Co-developed with Broadcom, announced October 2025
  • Designed to minimize prefill and communication phase delays

OpenAI's custom inference chip Jalapeño outperformed Nvidia's Blackwell system on SemiAnalysis' InferenceX benchmark, registering more tokens per user and more throughput per kilowatt, the company disclosed at Hot Chips on Tuesday. The comparison is notable, but the competitive window is narrow: Ho estimated Jalapeño would deploy "in very small volumes" at the end of 2026, with meaningful scale only in 2027 — by which point Nvidia's next-generation parts will likely be shipping. According to TechCrunch

The prefill and KV cache angle

The benchmark win is less about raw silicon and more about where inference bottlenecks actually live. Jalapeño was designed with Broadcom to minimize delays during prefill and communication phases, and to keep the KV cache local. "We designed Jalapeño to minimize data movement and communication delays," OpenAI said in a blog post, "so that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase." This is a direct response to the agentic inference workloads that have made KV cache management the central problem of the serving stack — a theme SemiAnalysis has pushed with its open-sourced $3M AgentX-InferenceXv3 dataset released just a day earlier.

The timing problem

The more uncomfortable question is whether beating today's Blackwell matters when the comparison target will be obsolete by the time Jalapeño ships at volume. Ho acknowledged the deployment timeline openly. OpenAI is also cutting API prices aggressively — GPT-5.6 Sol prices dropped 20-33% this week — which makes per-watt throughput a strategic lever, not just a technical one. The full-stack approach, with models, chips, and memory developed in concert, is the real structural advantage: it lets OpenAI attack inference phases that generic accelerators treat as uniform.

The company did not disclose absolute token counts, power draw, or die size, which limits the extent to which the InferenceX result can be independently verified. The company's blog post presents the results as a comparison against "currently available" state-of-the-art parts, a phrasing that leaves room for interpretation about what comes next.

Key Takeaways

  • OpenAI's Jalapeño chip beat Nvidia Blackwell on InferenceX at Hot Chips 2026, with more tokens per watt.
  • Volume deployment slips to 2027, raising questions about the competitive window.

What to watch

Watch for Nvidia's next-generation inference parts and whether Jalapeño's InferenceX advantage holds against them. Also track OpenAI's per-watt cost data in Q1 2027 as volume deployment begins, and whether the GPT-5.6 Sol price cuts reflect Jalapeño's efficiency or competitive pressure ahead of Anthropic's IPO.

OpenAI’s Jalapeño chip


Source: techcrunch.com

[Updated 25 Aug via the_decoder]

The chip's advantage extends beyond Blackwell: SemiAnalysis CEO Dylan Patel said Jalapeño also beats Nvidia's upcoming Rubin, noting "usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin." [per The Decoder] Specific figures from SemiAnalysis show Jalapeño delivers 1.5 to 1.9 times more AI work per watt at peak throughput, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times higher performance on highly interactive workloads. [per Next Big Future]


Sources cited in this article

  1. The Decoder
  2. Next Big Future
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 3 verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The Jalapeño result is less about winning today and more about the architectural bet OpenAI is making on where inference bottlenecks actually live. By explicitly targeting prefill and KV cache locality, OpenAI is acknowledging that agentic workloads — which drive long-context, multi-turn interactions — have different serving characteristics than the single-shot generation that dominated the early GPT era. This is a structural response to the problem SemiAnalysis has been documenting: the serving stack's center of gravity has shifted from compute throughput to memory bandwidth and data movement. The more interesting comparison is against the price cuts OpenAI announced this week. GPT-5.6 Sol dropped 20-33%, and Anthropic cut Claude Opus prices 50% — both moves framed as responses to enterprise cost pressure. If Jalapeño delivers the per-watt gains the benchmark suggests, OpenAI's cost structure improves just as the price war with Anthropic intensifies ahead of its IPO. The chip is not just a technical artifact; it is a margin lever in a competitive fight where API pricing is the primary weapon. The risk is timing. Blackwell is the comparison target today, but Nvidia's next-gen parts will likely ship before Jalapeño reaches meaningful volume in 2027. The benchmark is a snapshot, not a durable advantage. OpenAI's real bet is that the full-stack co-design — models, chips, memory, and networking developed together — compounds faster than Nvidia's general-purpose approach. That is a plausible thesis, but the evidence so far is one benchmark against a soon-to-be-obsolete target.
This story is part of
Hugging Face Becomes the Neutral Ground Where Google and Anthropic's Agent Protocol War Converges
As Claude Code's MCP dominance threatens Google Cloud, Hugging Face's unique position as partner to both players creates an unexpected convergence zone
Compare side-by-side
OpenAI vs Nvidia
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all