SWE-Bench Verified
benchmark→ stable
SWE-bench Verifiedswe-bench-verified
OpenAI-verified 500-issue subset of SWE-Bench. Approaching saturation in 2026 - most frontier models clear 80%+.
17Total Mentions
+0.05Sentiment (Neutral)
0.0%Velocity (7d)
First seen: Apr 25, 2026Last active: Jul 13, 2026
Signal Radar
Five-axis snapshot of this entity's footprint
Loading radar…
Mentions × Lab Attention
Weekly mentions (solid) and average article relevance (dotted)
mentionsrelevance
Loading timeline…
Timeline
No timeline events recorded yet.
Relationships
9Uses
Benchmarked On
Frequently appears with
10Entities that show up in the same articles — shared coverage, not a stated relationship.
Recent Articles
2Why Claude Code's 80.8% SWE-Bench Score and 1M Context Window Beat Codex
~Claude Code's 80.8% SWE-Bench score, 1M token context, and local execution make it the top choice for senior devs—use `claude code` in your terminal f
85 relevanceClaude Opus 4.8 Now Beats Gemini Pro 5 in Coding Benchmarks — What It
~Claude Opus 4.8 beats Gemini Pro 5 by 11 points on Fable 5. Claude Code users should run `claude code --model opus-4.8` for complex coding tasks.
100 relevance
Predictions
No predictions linked to this entity.
AI Discoveries
1- observationactive1d ago
Lifecycle: SWE-Bench Verified
SWE-Bench Verified is in 'declining' phase (0 mentions/3d, 1/14d, 17 total)
90% confidence
Sentiment History
6-W266-W29
Positive sentiment
Negative sentiment
Range: -1 to +1
| Week | Avg Sentiment | Mentions |
|---|---|---|
| 2026-W26 | 0.15 | 2 |
| 2026-W29 | 0.00 | 2 |