SWE-Bench Pro
Harder, contamination-resistant successor to SWE-Bench Verified: real GitHub issues with held-out tests. Where coding headroom remains.
Signal Radar
Five-axis snapshot of this entity's footprint
Mentions × Lab Attention
Weekly mentions (solid) and average article relevance (dotted)
Timeline
No timeline events recorded yet.
Relationships
5Benchmarked On
Uses
Frequently appears with
1Entities that show up in the same articles — shared coverage, not a stated relationship.
Recent Articles
3OpenAI GPT-5.6 Sol matches Fable 5 at 1/3 cost, adds multi-agent API
~OpenAI's GPT-5.6 Sol nearly matches Claude Fable 5 on aggregate benchmarks at one-third the cost, with new multi-agent and tool-calling APIs.
95 relevanceOpenAI Claims 54% Token Efficiency Gain on Agentic Coding in New Model
-OpenAI CEO Sam Altman claims 54% token efficiency gain on agentic coding for a new unnamed model, but no technical details or release date were provid
90 relevanceOpenAI Finds 30% of SWE-Bench Pro Tasks Are Broken, Pulls Endorsement
-OpenAI finds ~30% of SWE-Bench Pro tasks broken, pulls endorsement. Human reviewers flagged 249 flawed tasks.
95 relevance
Predictions
No predictions linked to this entity.
AI Discoveries
1- observationactiveJul 10, 2026
Velocity spike: SWE-Bench Pro
SWE-Bench Pro (benchmark) surged from 0 to 3 mentions in 3 days (new_surge).
80% confidence
Sentiment History
| Week | Avg Sentiment | Mentions |
|---|---|---|
| 2026-W23 | 0.00 | 1 |
| 2026-W28 | -0.33 | 3 |