SWE-Bench Verified
OpenAI-verified 500-issue subset of SWE-Bench. Approaching saturation in 2026 - most frontier models clear 80%+.
Signal Radar
Five-axis snapshot of this entity's footprint
Mentions × Lab Attention
Weekly mentions (solid) and average article relevance (dotted)
Timeline
No timeline events recorded yet.
Relationships
9Uses
Benchmarked On
Frequently appears with
10Entities that show up in the same articles — shared coverage, not a stated relationship.
Recent Articles
4Why Claude Code's 80.8% SWE-Bench Score and 1M Context Window Beat Codex
~Claude Code's 80.8% SWE-Bench score, 1M token context, and local execution make it the top choice for senior devs—use `claude code` in your terminal f
85 relevanceClaude Opus 4.8 Now Beats Gemini Pro 5 in Coding Benchmarks — What It
~Claude Opus 4.8 beats Gemini Pro 5 by 11 points on Fable 5. Claude Code users should run `claude code --model opus-4.8` for complex coding tasks.
100 relevanceClaude AI's $29 Kit Earns $0 in 12 Days — Kill-Criteria Clock Runs
~A Claude AI agent earned $0 in 12 days from a $29 kit, with 3 funnel visitors. A pre-written kill-criteria clock runs to July 3, 2026.
70 relevanceWorld Model MCP: Memory Layer That Cut SWE-bench Repeat Mistakes by +10.2 Points
~World Model MCP adds a temporal knowledge graph to Claude Code that learns from corrections, prevents repeated mistakes, and re-injects context after
95 relevance
Predictions
No predictions linked to this entity.
AI Discoveries
1- observationactive3d ago
Lifecycle: SWE-Bench Verified
SWE-Bench Verified is in 'active' phase (1 mentions/3d, 2/14d, 17 total)
90% confidence
Sentiment History
| Week | Avg Sentiment | Mentions |
|---|---|---|
| 2026-W26 | 0.15 | 2 |
| 2026-W29 | 0.00 | 2 |