gpt 5.6
30 articles about gpt 5.6 in AI news
OpenAI Ships GPT-5.6 Sol as Unified Reasoning Model
OpenAI unified ChatGPT reasoning under GPT-5.6 Sol for Plus/Pro, with Luna for free tiers. Internal eval shows 68% fewer factual errors.
OpenAI Cuts GPT-5.6 Luna Price 80% to $0.20/M Tokens
OpenAI cut GPT-5.6 Luna prices 80% to $0.20/M input tokens, citing Sol-optimized kernels that cut serving costs 20%. Luna now undercuts Gemini Flash-Lite and Claude Haiku.
GPT-5.6 Sol Leads DeepSWE at 72.7%, Beating Opus 5's 68.8%
GPT-5.6 Sol scores 72.7% on DeepSWE, beating Opus 5's 68.8%. The undocumented benchmark tests autonomous SWE agents.
Gemini 3.6 Flash Hits 83% on Computer Use, Beats GPT-5.6 and Grok
Gemini 3.6 Flash scored 83.0% on OSWorld-Verified, beating GPT-5.6 and Grok at $7.50 per million tokens, breaking the cheap-tier capability trade-off.
GPT-5.6 Sol on Cerebras Hits 750 Token/s
GPT-5.6 Sol on Cerebras claimed at 750 token/s, but no official data or model release exists. Unverified claim needs vendor confirmation.
OpenAI GPT-5.6 Sol, Terra, Luna Launch on Bedrock at Same Price
OpenAI's GPT-5.6 Sol, Terra, and Luna launch on Amazon Bedrock at matching first-party pricing. Sol scores 80 on Coding Agent Index.
OpenAI GPT-5.6 Sol matches Fable 5 at 1/3 cost, adds multi-agent API
OpenAI's GPT-5.6 Sol nearly matches Claude Fable 5 on aggregate benchmarks at one-third the cost, with new multi-agent and tool-calling APIs.
OpenAI GPT-5.6 Launches Thursday After US Gov't Lifts Ban
OpenAI's GPT-5.6 Sol launches Thursday after US gov't lifts ban. It beats Claude Mythos 5 on benchmarks at half the cost.
GPT-5.6 Sol, Terra, Luna: Benchmark Performance Depends on Which Test You Use
OpenAI released GPT-5.6 as three tiers—Sol, Terra, Luna—on June 27, 2026. Sol tops Terminal-Bench 2.1 but trails competitors on other benchmarks. The release shifts focus to tiered pricing and efficiency, but access remains restricted.
OpenAI Launches GPT-5.6 Sol Under US Government Restrictions
OpenAI's GPT-5.6 Sol beats Claude Mythos 5 in agentic coding (88.8% vs 88%) but US government restricts access to select partners, a policy OpenAI calls unsustainable.
White House Orders OpenAI to Gate GPT-5.6 Release per Customer
White House orders OpenAI to gate GPT-5.6 release per customer, mirroring Anthropic's voluntary suspension of Claude Mythos under regulatory pressure.
GPT-Red: OpenAI's LLM Super-Hacker Finds 84% of Attacks, Humans 13%
OpenAI's GPT-Red LLM finds 84% of attacks vs 13% for humans, hardening GPT-5.6 Sol. Automated red-teaming shifts safety paradigm.
GPT-4 Held Top Spot 52 Weeks; Today's Models Last 7
GPT-4 dominated the ECI for a year. Today's top models last 7 weeks median, with 17 leadership changes since Feb 2024.
Multi-User LLM Agents Struggle: Gemini 3 Pro Scores 85.6% on Muses-Bench
A new benchmark reveals LLMs struggle with multi-user scenarios where agents face conflicting instructions. Gemini 3 Pro leads but only achieves 85.6% average, with privacy-utility tradeoffs proving particularly difficult.
DeepSeek V4 Flash 0731 Hits 50 on Intelligence Index at $0.14/M Tokens
DeepSeek V4 Flash 0731 scores 50 on Intelligence Index, one point behind GPT-5.6 Luna at ~60% lower cost. 304B params, $0.14/M input pricing.
OpenAI hits 38.3% on ARC-AGI-3 with custom API, bypassing official harness
OpenAI's GPT-5.6 Sol scored 38.3% on ARC-AGI-3 with custom API settings, beating Opus 5's 30.2%, but scored 7.8% in the official harness, exposing benchmark parity issues.
Microsoft MAI-Cyber-1-Flash Hits 96% on CyberGym
Microsoft's MAI-Cyber-1-Flash scores 96% on CyberGym, cutting costs 50% by handling 90% of security tasks locally while routing complex cases to GPT-5.4.
METR's 'Expenditure Horizon': AI Agents Break Even at $3,300
METR's expenditure horizon metric shows AI agents break even at $0–$3,300 on NanoGPT, vs $2,500 per 1% speedup for humans. GPT-5 and Opus-4.1 pro lead, but blind spots remain.
Port Claude Code Workflows to Codex
gpt-workflow brings Claude Code-style deterministic workflows to Codex CLI with resumable journals and JSON schema validation. Install via Codex plugin and store workflows under .codex/workflows/.
DeepSeek seeks fresh $71B round weeks after $7B close
DeepSeek seeks $71B valuation round weeks after $7B close. Capital for data centers and custom chips to sustain 11x cheaper pricing than GPT-5.5.
New Research Improves Agentic RAG Efficiency with Contextualization and De-duplication Modules
Researchers propose test-time modifications to agentic RAG systems, adding contextualization and de-duplication modules. Their best variant achieves 5.6% higher accuracy and 10.5% fewer retrieval turns, making complex question-answering more efficient.
Token-Saving Tools Overpromise: Real Benchmark Shows 6–32% Savings, Not 60–90%
Token-saving tools deliver 6–32% savings, not 60–90%. In Claude Code, lazy MCP loading means tools often go unused—enable them with hooks and measure full sessions.
Claude Tool Use: Fable 5 Beats Opus 4.8 at 1.00 Calls
SemiAnalysis analyzed 2.27M Claude responses, finding Fable 5 averages 1.00 tool calls per response versus 0.76 for Opus 4.8. The Opus line shows a downward trend.
Opus 5 Generates Playable Game for $423 in One Prompt
Opus 5 generated a playable game from one prompt, using 690M tokens at $423, per @kimmonismus. One person replaced a dev team.
OpenAI's Astra Solves 10 Open Math Problems, Costs $2K
OpenAI's Astra solved ten open math problems at ~$2K token cost, formalized in Lean. First model to face U.S. government review.
Nvidia-OpenAI Talks Hit $250B; NVL72 Ships 96GB HBM3e
Nvidia in talks to invest $250B in OpenAI, six times Microsoft's stake. NVL72 now ships 96GB HBM3e; $25B bond planned.
Moonshot AI Releases 1.56T-Parameter Kimi K3, Requires 2x B200 Nodes
Moonshot AI released Kimi K3, a 1.56T parameter MoE model at 1561 GB, requiring 2x B200 nodes. No benchmarks disclosed.
Anthropic Ships Claude Opus 5: Fable-Level Intelligence at Half the Price
Anthropic released Claude Opus 5 on July 24 with a 1M token context, 128k output, and Fable-5-approaching intelligence at half the price, unchanged from Opus 4.8.
Opus 5 Hits 0% Prompt Injection Rate in Browser Agents
Anthropic's Opus 5 with Auto Mode achieved 0% prompt injection success across 129 tests, challenging OpenAI's view that the problem is unsolvable.
Jensen Huang: DeepSeek, Kimi open models boost Nvidia sales
Jensen Huang says Chinese open models DeepSeek and Kimi boost Nvidia GPU demand, not threaten it. Market misunderstood their impact twice.