OpenAI cut gpt-5-6-luna" class="entity-chip">GPT-5.6 Luna prices by 80% to $0.20 per million input tokens starting July 30. The move puts Luna below Google's Gemini 3.1 Flash-Lite and one-fifth of Anthropic's Claude Haiku 4.5 input price.
Key facts
- GPT-5.6 Luna input price: $0.20/M tokens, down 80%
- GPT-5.6 Terra cut 20%, effective July 30
- Sol's kernel rewrites cut serving costs 20%
- Luna now 1/5th Claude Haiku 4.5 input price
- DeepSeek V4 Flash scores 50, one point behind Luna
OpenAI's price announcement marks a structural shift in the low-cost tier. GPT-5.6 Luna drops from parity with Claude Haiku 4.5 at $1/$5 to $0.20/$1.20 per million tokens, while GPT-5.6 Terra gets a modest 20% reduction. The company frames the cuts as a byproduct of internal efficiency gains, not a competitive response — but the timing is telling.
How Sol paid for the cuts
OpenAI credits GPT-5.6 Sol, its top-tier model, with enabling the price reductions. In the companion post "How GPT-5.6 fuses frontier intelligence with frontier efficiency", the company describes using Sol to optimize load balancing and, more significantly, to rewrite production kernels.
We also used GPT‑5.6 Sol to optimize the model's forward pass... GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels.
This worked because GPT-5.6 was trained to write and improve kernels in Triton and Gluon, two open-source GPU programming languages OpenAI maintains. The combined kernel and load-balancing work reduced end-to-end serving costs by 20%, according to the company's blog post.
The competitive floor just moved
The Luna price reset changes the economics for high-volume agentic workloads. At $0.20/$1.20, Luna undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's input price — previously it cost the same. Simon Willison switched his agent.datasette.io demo from Gemini 3.1 Flash-Lite to Luna the same day.
The pressure isn't one-directional. DeepSeek's V4 Flash "0731" update jumped ten points to 50 on the Artificial Analysis Intelligence Index, The Decoder reports — one point behind GPT-5.6 Luna at roughly 60% lower cost per task. Microsoft's MAI models and cheap Chinese providers likely forced OpenAI's hand, per The Decoder.
The real story is that OpenAI has turned its frontier model into a cost-reduction tool. Sol isn't just the most capable model — it's now an internal optimization engine that rewrites the kernels that serve cheaper models. That's a moat DeepSeek and Google can't easily replicate, because it requires frontier-level capability applied to infrastructure itself. The 80% cut isn't a discount; it's a demonstration that OpenAI's serving stack now benefits from its own best model.
Key Takeaways
- OpenAI cut GPT-5.6 Luna prices 80% to $0.20/M input tokens, citing Sol-optimized kernels that cut serving costs 20%.
- Luna now undercuts Gemini Flash-Lite and Claude Haiku.
What to watch
Watch whether Anthropic and Google respond with matching cuts to Haiku and Flash-Lite within the next 30 days. Also track DeepSeek's V4 Flash pricing — at 60% lower cost per task, it may force OpenAI to cut Luna again before Q4 earnings.
Source: openai.com








