GPT-5
GPT-5 is OpenAI's fifth-generation multimodal large language model, released on February 16, 2026. It accepts text and image inputs and generates text, succeeding GPT-4. Upon release, it achieved an Arena Elo score of 1450 on the LMSYS Chatbot Arena leaderboard and a SWE-bench Verified score of 80.0%, according to OpenAI's published technical report. By comparison, GPT-4 Turbo posted an Arena Elo of 1250 and a SWE-bench Verified score of 67.0% in the same evaluation framework. API pricing was set at $1.75 per million input tokens and $14.00 per million output tokens. GPT-5 matters now because its February 2026 debut marks a measurable leap in reasoning and coding capabilities, establishing a new empirical baseline in public, third-party-evaluated benchmarks and directly shaping the competitive landscape against contemporaneous frontier models from Anthropic and Google DeepMind.
OpenAI's GPT-5, a multimodal LLM released February 2026, posts an Arena Elo of 1450 and a SWE-bench Verified score of 80.0. It deploys Sparse Mixture of Experts and Chain-of-Thought prompting, but the graph reveals a crowded competitive field: it is locked in direct competition with Claude Opus 4.7, MolmoAct2, and Claude Opus 4.6. The incoming edge from Anthropic and Google as developers of rival models signals ecosystem pressure. Recent headlines show its successor, GPT-5.6 Sol, already leading DeepSWE at 72.7%, while CAS ZhiJing beats GPT-5.5 on social cognition. Despite robust mention volume (15 in 30 days), GPT-5 is being flanked from above by its own lineage and from below by specialized challengers. The central tendency bias research it uses may limit its output distinctiveness.
- ·Achieves Arena Elo 1450 and SWE-bench 80.0 at launch
- ·Competes with Claude Opus 4.7, MolmoAct2, and Claude Opus 4.6
- ·Deploys Sparse MoE and Chain-of-Thought techniques
- ·Successor GPT-5.6 Sol already outperforms on DeepSWE (72.7%)
- ·Faces pressure from Anthropic and Google as developer rivals
Signal Radar
Five-axis snapshot of this entity's footprint
Mentions × Lab Attention
Weekly mentions (solid) and average article relevance (dotted)
Timeline
4- Research MilestoneMay 11, 2026
Study reveals GPT-5 exhibits central tendency bias in clinical scoring, compressing predictions toward scale midpoint.
View source- mae zero shot:
- 0.67
- within 1 accuracy:
- 92%
- Research MilestoneMar 6, 2026
Evaluation study published on arXiv assessing its clinical reasoning capabilities
View source - Product LaunchMay 31, 2025
Reportedly became available according to user social media post
View source
Relationships
14Developed
Deploys
Competes With
Uses
Frequently appears with
10Entities that show up in the same articles — shared coverage, not a stated relationship.
Recent Articles
15OpenAI hits 38.3% on ARC-AGI-3 with custom API, bypassing official harness
~OpenAI's GPT-5.6 Sol scored 38.3% on ARC-AGI-3 with custom API settings, beating Opus 5's 30.2%, but scored 7.8% in the official harness, exposing ben
100 relevanceMicrosoft MAI-Cyber-1-Flash Hits 96% on CyberGym
~Microsoft's MAI-Cyber-1-Flash scores 96% on CyberGym, cutting costs 50% by handling 90% of security tasks locally while routing complex cases to GPT-5
100 relevanceMETR's 'Expenditure Horizon': AI Agents Break Even at $3,300
+METR's expenditure horizon metric shows AI agents break even at $0–$3,300 on NanoGPT, vs $2,500 per 1% speedup for humans. GPT-5 and Opus-4.1 pro lead
90 relevanceCAS ZhiJing Beats GPT-5.5 on Social Cognition with FLARE Training
~CAS ICT releases ZhiJing social intelligence system, Zing model beats GPT-5.5 on social cognition via FLARE training.
88 relevanceEpoch AI: Google's Colossus 1 Training Compute Hits 1e26 FLOP
~Google's Colossus 1 used 1e26 FLOP at $4.6B, per Epoch AI. It is the largest known training run, signaling a new capital scale.
100 relevanceGPT-5.6 Sol Leads DeepSWE at 72.7%, Beating Opus 5's 68.8%
~GPT-5.6 Sol scores 72.7% on DeepSWE, beating Opus 5's 68.8%. The undocumented benchmark tests autonomous SWE agents.
100 relevanceBenchmark lets image models answer in pixels, not text
~New 'Show, Don't Tell' benchmark tests spatial cognition via pixel-level outputs. GPT Image 2 solves 37% of cases missed by GPT-5.4, highlighting a ga
85 relevanceGemini 3.6 Flash Hits 83% on Computer Use, Beats GPT-5.6 and Grok
~Gemini 3.6 Flash scored 83.0% on OSWorld-Verified, beating GPT-5.6 and Grok at $7.50 per million tokens, breaking the cheap-tier capability trade-off.
87 relevanceTraders Bet Claude Opus 4.8 Launch Imminent as Options Spike
~Traders bet Anthropic will launch Claude Opus 4.8 within days, based on options market activity. The model would succeed Opus 4.7 (69.2% SWE-bench Pro
75 relevanceGPT-5.6 Sol on Cerebras Hits 750 Token/s
~GPT-5.6 Sol on Cerebras claimed at 750 token/s, but no official data or model release exists. Unverified claim needs vendor confirmation.
97 relevanceDeepSeek seeks fresh $71B round weeks after $7B close
~DeepSeek seeks $71B valuation round weeks after $7B close. Capital for data centers and custom chips to sustain 11x cheaper pricing than GPT-5.5.
100 relevanceChatGPT returns to WhatsApp in EU after Meta forced to open platform
~OpenAI re-enabled ChatGPT on WhatsApp in the EEA after EU forced Meta to open its platform. Users reach GPT-5.5 via 1-800-CHATGPT.
93 relevanceOpenAI GPT-5.6 Sol, Terra, Luna Launch on Bedrock at Same Price
~OpenAI's GPT-5.6 Sol, Terra, and Luna launch on Amazon Bedrock at matching first-party pricing. Sol scores 80 on Coding Agent Index.
100 relevanceOpenAI GPT-5.6 Sol matches Fable 5 at 1/3 cost, adds multi-agent API
~OpenAI's GPT-5.6 Sol nearly matches Claude Fable 5 on aggregate benchmarks at one-third the cost, with new multi-agent and tool-calling APIs.
95 relevanceOpenAI GPT-5.6 Launches Thursday After US Gov't Lifts Ban
~OpenAI's GPT-5.6 Sol launches Thursday after US gov't lifts ban. It beats Claude Mythos 5 on benchmarks at half the cost.
100 relevance
Predictions
No predictions linked to this entity.
AI Discoveries
6- observationactive8h ago
Novel co-occurrence: Opus 5 + GPT-5
Opus 5 (ai_model) and GPT-5 (ai_model) appeared together in 2 articles this week but have NEVER co-occurred before and have no existing relationship. This is a potential breaking story signal.
85% confidence - observationactive2d ago
Novel co-occurrence: DeepSWE + GPT-5
DeepSWE (technology) and GPT-5 (ai_model) appeared together in 2 articles this week but have NEVER co-occurred before and have no existing relationship. This is a potential breaking story signal.
85% confidence - observationactive3d ago
Velocity spike: GPT-5
GPT-5 (ai_model) surged from 1 to 4 mentions in 3 days (velocity_spike).
80% confidence - discoveryactiveJul 21, 2026
Causal: Anthropic launches Claude Code with MCP → OpenAI will acquire or build an MCP-comp
Cause: Anthropic launches Claude Code with MCP (22 mentions/7d, 124 shared with MCP) Effect: OpenAI responds with GPT-5 (4 mentions/7d) but lacks equivalent protocol standardization Predicted next: OpenAI will acquire or build an MCP-compatible agent framework within 60 days, or risk Claude Code bec
75% confidence - observationactiveJul 17, 2026
Lifecycle: GPT-5
GPT-5 is in 'established' phase (2 mentions/3d, 5/14d, 42 total)
90% confidence - discoveryactiveJul 16, 2026
The 'Agent Wars' Are Actually an Infrastructure Play: Anthropic vs. OpenAI via MCP vs. Function Calling
Claude Code (23 mentions/7d) and GPT-5 (4 mentions/7d) are unconnected pairs, but Anthropic & OpenAI have 265 shared articles. This is non-obvious because the surface narrative is about model quality, but the real battle is over agent infrastructure standards. Anthropic is betting on MCP as the univ
88% confidence
Sentiment History
| Week | Avg Sentiment | Mentions |
|---|---|---|
| 2026-W25 | 0.10 | 2 |
| 2026-W26 | 0.00 | 5 |
| 2026-W27 | -0.20 | 1 |
| 2026-W28 | 0.00 | 2 |
| 2026-W29 | 0.00 | 4 |
| 2026-W30 | 0.02 | 5 |
| 2026-W31 | 0.13 | 4 |