Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

gpt 5.6

30 articles about gpt 5.6 in AI news

OpenAI Ships GPT-5.6 Sol as Unified Reasoning Model

OpenAI unified ChatGPT reasoning under GPT-5.6 Sol for Plus/Pro, with Luna for free tiers. Internal eval shows 68% fewer factual errors.

100% relevant

OpenAI Cuts GPT-5.6 Luna Price 80% to $0.20/M Tokens

OpenAI cut GPT-5.6 Luna prices 80% to $0.20/M input tokens, citing Sol-optimized kernels that cut serving costs 20%. Luna now undercuts Gemini Flash-Lite and Claude Haiku.

100% relevant

GPT-5.6 Sol Leads DeepSWE at 72.7%, Beating Opus 5's 68.8%

GPT-5.6 Sol scores 72.7% on DeepSWE, beating Opus 5's 68.8%. The undocumented benchmark tests autonomous SWE agents.

100% relevant

Gemini 3.6 Flash Hits 83% on Computer Use, Beats GPT-5.6 and Grok

Gemini 3.6 Flash scored 83.0% on OSWorld-Verified, beating GPT-5.6 and Grok at $7.50 per million tokens, breaking the cheap-tier capability trade-off.

87% relevant

GPT-5.6 Sol on Cerebras Hits 750 Token/s

GPT-5.6 Sol on Cerebras claimed at 750 token/s, but no official data or model release exists. Unverified claim needs vendor confirmation.

97% relevant

OpenAI GPT-5.6 Sol, Terra, Luna Launch on Bedrock at Same Price

OpenAI's GPT-5.6 Sol, Terra, and Luna launch on Amazon Bedrock at matching first-party pricing. Sol scores 80 on Coding Agent Index.

100% relevant

OpenAI GPT-5.6 Sol matches Fable 5 at 1/3 cost, adds multi-agent API

OpenAI's GPT-5.6 Sol nearly matches Claude Fable 5 on aggregate benchmarks at one-third the cost, with new multi-agent and tool-calling APIs.

95% relevant

OpenAI GPT-5.6 Launches Thursday After US Gov't Lifts Ban

OpenAI's GPT-5.6 Sol launches Thursday after US gov't lifts ban. It beats Claude Mythos 5 on benchmarks at half the cost.

100% relevant

GPT-5.6 Sol, Terra, Luna: Benchmark Performance Depends on Which Test You Use

OpenAI released GPT-5.6 as three tiers—Sol, Terra, Luna—on June 27, 2026. Sol tops Terminal-Bench 2.1 but trails competitors on other benchmarks. The release shifts focus to tiered pricing and efficiency, but access remains restricted.

76% relevant

OpenAI Launches GPT-5.6 Sol Under US Government Restrictions

OpenAI's GPT-5.6 Sol beats Claude Mythos 5 in agentic coding (88.8% vs 88%) but US government restricts access to select partners, a policy OpenAI calls unsustainable.

100% relevant

White House Orders OpenAI to Gate GPT-5.6 Release per Customer

White House orders OpenAI to gate GPT-5.6 release per customer, mirroring Anthropic's voluntary suspension of Claude Mythos under regulatory pressure.

100% relevant

GPT-Red: OpenAI's LLM Super-Hacker Finds 84% of Attacks, Humans 13%

OpenAI's GPT-Red LLM finds 84% of attacks vs 13% for humans, hardening GPT-5.6 Sol. Automated red-teaming shifts safety paradigm.

91% relevant

GPT-4 Held Top Spot 52 Weeks; Today's Models Last 7

GPT-4 dominated the ECI for a year. Today's top models last 7 weeks median, with 17 leadership changes since Feb 2024.

84% relevant

Multi-User LLM Agents Struggle: Gemini 3 Pro Scores 85.6% on Muses-Bench

A new benchmark reveals LLMs struggle with multi-user scenarios where agents face conflicting instructions. Gemini 3 Pro leads but only achieves 85.6% average, with privacy-utility tradeoffs proving particularly difficult.

92% relevant

DeepSeek V4 Flash 0731 Hits 50 on Intelligence Index at $0.14/M Tokens

DeepSeek V4 Flash 0731 scores 50 on Intelligence Index, one point behind GPT-5.6 Luna at ~60% lower cost. 304B params, $0.14/M input pricing.

100% relevant

OpenAI hits 38.3% on ARC-AGI-3 with custom API, bypassing official harness

OpenAI's GPT-5.6 Sol scored 38.3% on ARC-AGI-3 with custom API settings, beating Opus 5's 30.2%, but scored 7.8% in the official harness, exposing benchmark parity issues.

100% relevant

Microsoft MAI-Cyber-1-Flash Hits 96% on CyberGym

Microsoft's MAI-Cyber-1-Flash scores 96% on CyberGym, cutting costs 50% by handling 90% of security tasks locally while routing complex cases to GPT-5.4.

100% relevant

METR's 'Expenditure Horizon': AI Agents Break Even at $3,300

METR's expenditure horizon metric shows AI agents break even at $0–$3,300 on NanoGPT, vs $2,500 per 1% speedup for humans. GPT-5 and Opus-4.1 pro lead, but blind spots remain.

90% relevant

Port Claude Code Workflows to Codex

gpt-workflow brings Claude Code-style deterministic workflows to Codex CLI with resumable journals and JSON schema validation. Install via Codex plugin and store workflows under .codex/workflows/.

90% relevant

DeepSeek seeks fresh $71B round weeks after $7B close

DeepSeek seeks $71B valuation round weeks after $7B close. Capital for data centers and custom chips to sustain 11x cheaper pricing than GPT-5.5.

100% relevant

New Research Improves Agentic RAG Efficiency with Contextualization and De-duplication Modules

Researchers propose test-time modifications to agentic RAG systems, adding contextualization and de-duplication modules. Their best variant achieves 5.6% higher accuracy and 10.5% fewer retrieval turns, making complex question-answering more efficient.

99% relevant

Token-Saving Tools Overpromise: Real Benchmark Shows 6–32% Savings, Not 60–90%

Token-saving tools deliver 6–32% savings, not 60–90%. In Claude Code, lazy MCP loading means tools often go unused—enable them with hooks and measure full sessions.

90% relevant

Claude Tool Use: Fable 5 Beats Opus 4.8 at 1.00 Calls

SemiAnalysis analyzed 2.27M Claude responses, finding Fable 5 averages 1.00 tool calls per response versus 0.76 for Opus 4.8. The Opus line shows a downward trend.

100% relevant

Opus 5 Generates Playable Game for $423 in One Prompt

Opus 5 generated a playable game from one prompt, using 690M tokens at $423, per @kimmonismus. One person replaced a dev team.

100% relevant

OpenAI's Astra Solves 10 Open Math Problems, Costs $2K

OpenAI's Astra solved ten open math problems at ~$2K token cost, formalized in Lean. First model to face U.S. government review.

92% relevant

Nvidia-OpenAI Talks Hit $250B; NVL72 Ships 96GB HBM3e

Nvidia in talks to invest $250B in OpenAI, six times Microsoft's stake. NVL72 now ships 96GB HBM3e; $25B bond planned.

100% relevant

Moonshot AI Releases 1.56T-Parameter Kimi K3, Requires 2x B200 Nodes

Moonshot AI released Kimi K3, a 1.56T parameter MoE model at 1561 GB, requiring 2x B200 nodes. No benchmarks disclosed.

100% relevant

Anthropic Ships Claude Opus 5: Fable-Level Intelligence at Half the Price

Anthropic released Claude Opus 5 on July 24 with a 1M token context, 128k output, and Fable-5-approaching intelligence at half the price, unchanged from Opus 4.8.

100% relevant

Opus 5 Hits 0% Prompt Injection Rate in Browser Agents

Anthropic's Opus 5 with Auto Mode achieved 0% prompt injection success across 129 tests, challenging OpenAI's view that the problem is unsolvable.

100% relevant

Jensen Huang: DeepSeek, Kimi open models boost Nvidia sales

Jensen Huang says Chinese open models DeepSeek and Kimi boost Nvidia GPU demand, not threaten it. Market misunderstood their impact twice.

99% relevant