Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

OpenAI CEO on stage at a launch event, with a large screen displaying a price chart and the GPT-5.6 Luna logo…
Products & LaunchesBreakthroughScore: 100

OpenAI Cuts GPT-5.6 Luna Price 80% to $0.20/M Tokens

OpenAI cut GPT-5.6 Luna prices 80% to $0.20/M input tokens, citing Sol-optimized kernels that cut serving costs 20%. Luna now undercuts Gemini Flash-Lite and Claude Haiku.

·Jul 30, 2026·4 min read··148 views·AI-Generated·Report error
Share:
Source: openai.comvia openai_blog, the_decoder, simon_willison, scmp_tech, @teortaxestexWidely Reported
How much did OpenAI cut GPT-5.6 Luna and Terra prices?

OpenAI cut GPT-5.6 Luna prices by 80% to $0.20 per million input tokens and $1.20 per million output tokens, effective July 30, 2026. GPT-5.6 Terra dropped 20%. The cuts follow GPT-5.6 Sol optimizing inference kernels and load balancing, reducing end-to-end serving costs by 20%.

TL;DR

GPT-5.6 Luna input price drops 80% to $0.20 per million tokens · OpenAI credits GPT-5.6 Sol for 20% serving cost cut · Luna now undercuts Gemini 3.1 Flash-Lite and Claude Haiku 4.5

OpenAI cut gpt-5-6-luna" class="entity-chip">GPT-5.6 Luna prices by 80% to $0.20 per million input tokens starting July 30. The move puts Luna below Google's Gemini 3.1 Flash-Lite and one-fifth of Anthropic's Claude Haiku 4.5 input price.

Key facts

  • GPT-5.6 Luna input price: $0.20/M tokens, down 80%
  • GPT-5.6 Terra cut 20%, effective July 30
  • Sol's kernel rewrites cut serving costs 20%
  • Luna now 1/5th Claude Haiku 4.5 input price
  • DeepSeek V4 Flash scores 50, one point behind Luna

OpenAI's price announcement marks a structural shift in the low-cost tier. GPT-5.6 Luna drops from parity with Claude Haiku 4.5 at $1/$5 to $0.20/$1.20 per million tokens, while GPT-5.6 Terra gets a modest 20% reduction. The company frames the cuts as a byproduct of internal efficiency gains, not a competitive response — but the timing is telling.

How Sol paid for the cuts

OpenAI credits GPT-5.6 Sol, its top-tier model, with enabling the price reductions. In the companion post "How GPT-5.6 fuses frontier intelligence with frontier efficiency", the company describes using Sol to optimize load balancing and, more significantly, to rewrite production kernels.

We also used GPT‑5.6 Sol to optimize the model's forward pass... GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels.

This worked because GPT-5.6 was trained to write and improve kernels in Triton and Gluon, two open-source GPU programming languages OpenAI maintains. The combined kernel and load-balancing work reduced end-to-end serving costs by 20%, according to the company's blog post.

The competitive floor just moved

The Luna price reset changes the economics for high-volume agentic workloads. At $0.20/$1.20, Luna undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's input price — previously it cost the same. Simon Willison switched his agent.datasette.io demo from Gemini 3.1 Flash-Lite to Luna the same day.

The pressure isn't one-directional. DeepSeek's V4 Flash "0731" update jumped ten points to 50 on the Artificial Analysis Intelligence Index, The Decoder reports — one point behind GPT-5.6 Luna at roughly 60% lower cost per task. Microsoft's MAI models and cheap Chinese providers likely forced OpenAI's hand, per The Decoder.

The real story is that OpenAI has turned its frontier model into a cost-reduction tool. Sol isn't just the most capable model — it's now an internal optimization engine that rewrites the kernels that serve cheaper models. That's a moat DeepSeek and Google can't easily replicate, because it requires frontier-level capability applied to infrastructure itself. The 80% cut isn't a discount; it's a demonstration that OpenAI's serving stack now benefits from its own best model.

Key Takeaways

  • OpenAI cut GPT-5.6 Luna prices 80% to $0.20/M input tokens, citing Sol-optimized kernels that cut serving costs 20%.
  • Luna now undercuts Gemini Flash-Lite and Claude Haiku.

What to watch

OpenAI Slashes Prices on GPT-5.6 Luna and Terra Models / X

Watch whether Anthropic and Google respond with matching cuts to Haiku and Flash-Lite within the next 30 days. Also track DeepSeek's V4 Flash pricing — at 60% lower cost per task, it may force OpenAI to cut Luna again before Q4 earnings.


Source: openai.com

[Updated 01 Aug via scmp_tech]

Sam Altman personally announced the price cut on X, framing it as a defensive move against Chinese rivals, [per SCMP]. The same day, DeepSeek released V4 Flash 0731, a 304-billion-parameter model (167GB on Hugging Face) with enhanced agentic capabilities, priced at $0.14/M input and $0.27/M output — roughly half Luna's input cost and 60% cheaper per task. Artificial Analysis ranks it ahead of MiniMax M3 (428B) and shows it outperforming models costing ten times more per task, positioning it as the best value-per-intelligence option currently available.

[Updated 01 Aug via scmp_tech]

OpenAI is reportedly developing a new model family, codenamed "Astra," designed to enable multiple agents to collaborate on complex problems over extended periods, according to The Decoder. CEO Sam Altman has already demoed Astra to policymakers in Washington, though OpenAI has not decided whether to release it as GPT-6 or a GPT-5 variant. The announcement was made by dropping ten previously unsolved math solutions, signaling a major leap in frontier capability. This development adds a new dimension to the competitive landscape, as OpenAI's frontier model is not only driving cost reductions but also paving the way for next-generation agentic workloads that could redefine how AI systems tackle long-horizon tasks.

[Updated 02 Aug via scmp_tech]

SCMP reports that OpenAI also introduced a "fast mode" for its flagship GPT-5.6 Sol, boosting performance speed up to 2.5 times, alongside the price cuts. The same article notes that the price reductions instantly placed Luna in the "most attractive" tier for intelligence-per-dollar, ahead of Zhipu AI's GLM-5.2 and MiniMax's M3, according to Artificial Analysis.


Sources cited in this article

  1. The Decoder
  2. SCMP
  3. The Decoder. CEO Sam Altman
  4. Artificial Analysis.
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 4 verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The 80% cut is less a price war move than a structural proof point. OpenAI has closed the loop where its most capable model optimizes the serving stack for its cheapest models. This is a compounding advantage: every Sol improvement in kernel efficiency drops costs across the entire GPT-5.6 family, not just the flagship. Competitors like DeepSeek are optimizing for benchmark score per dollar; OpenAI is optimizing the cost curve itself. The DeepSeek V4 Flash result complicates the narrative. At 50 on the Artificial Analysis Intelligence Index, it's one point behind Luna at 60% lower cost. That's a direct challenge to OpenAI's value proposition in the low tier. If DeepSeek maintains this price-performance ratio, OpenAI's 80% cut may not be enough — it could be the first of several. The Triton/Gluon angle is worth noting. OpenAI maintaining its own GPU programming languages gives it a kernel-optimization flywheel that closed-source competitors like Anthropic lack. Google has XLA, but it's not designed for frontier-model self-optimization. This is a structural moat that won't show up in benchmark tables.
This story is part of
The Protocol Schism: Anthropic's MCP Stack vs. OpenAI's Agent Lock-In
How a developer convention is splitting AI into two incompatible ecosystems, with Meta and Google caught in the middle
Compare side-by-side
OpenAI vs Anthropic
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all