benchmark
30 articles about benchmark in AI news
Hugging Face Roundup: 10 Papers Push Harder Benchmarks, Multi-Agent AI
HF's weekly list of 10 papers signals a shift to harder benchmarks and multi-agent AI. Includes SWE-Bench ProMax and Alaya-EVOKE.
SPOT Distillation Beats OPD, EOPD on Reasoning Benchmarks
SPOT distillation beats OPD and EOPD on reasoning benchmarks by sparsely probing key positions and using outcome-calibrated targets. Reported by @HuggingPapers.
StreamArena: 243-Video Benchmark for Hour-Long Streaming AI
StreamArena launches with 243 hour-long videos and 3,646 open-ended tasks for interactive streaming video understanding, pushing beyond short-clip benchmarks. No baselines yet.
Agent Memory Benchmarks Get a Shared 5,000-Question Test
A 20+ institution consortium launched a shared agent memory benchmark with 5,000 questions and one answer model, targeting the attribution problem in memory startup claims. First rankings due mid-August.
Supabase's Evals Benchmark Just Gave Claude Code a Real-World Report Card
Supabase Evals is an open-source benchmark that scores Claude Code, Codex, and OpenCode on real Supabase tasks. Run `supabase eval` on your repo to find agent weaknesses.
ClBench-V: New Benchmark Tests Multimodal Contextual Learning in 3 Dimensions
ClBench-V benchmark from @HuggingPapers tests multimodal contextual learning across three dimensions: grounding, application, and knowledge learning. No results disclosed yet.
Ant Ling-3.0-flash Beats 1T-Ring-2.6 on 11 of 12 Benchmarks
Ant's 124B-param Ling-3.0-flash with 5.1B activated beats 1T-Ring-2.6 in 11 of 12 benchmarks, tying DeepSeek V4 Flash. Sparse activation economics are the story.
FutureX Refactoring Benchmark: 40% Faster Than Claude Code, 80% Test Pass Rate
FutureX refactored code 40% faster than Claude Code in a controlled benchmark, with an 80% initial test pass rate vs 60%. The specialized agent required 4 minutes of review per task versus 7 minutes for Claude Code.
Benchmark lets image models answer in pixels, not text
New 'Show, Don't Tell' benchmark tests spatial cognition via pixel-level outputs. GPT Image 2 solves 37% of cases missed by GPT-5.4, highlighting a gap in text-based spatial reasoning.
ActiveVision Benchmark: Humans 96.1%, Best AI 10.6%
ActiveVision benchmark: humans 96.1%, best AI 10.6%. The 85.5-point gap reveals fundamental limits in iterative visual reasoning for current models.
90 Hours of Black Myth: Wukong Fuel New World Model Benchmark
A new survey and benchmark rethinks interactive world models as game engines, with a data engine collecting over 90 hours of Black Myth: Wukong gameplay.
KeyFrame-Compass Benchmark Targets Keyframe Video Generation Gaps
KeyFrame-Compass is the first benchmark for keyframe-conditioned video generation, with 386 samples and six metrics.
gdb: Benchmarks Saturate Too Fast for Reliable AI Progress Tracking
@gdb notes benchmarks saturate quickly. This undermines AI progress tracking and may force shift to dynamic evaluations.
Soofi S 30B-A3B: German open model tops English, German benchmarks
German consortium releases Soofi S 30B-A3B, an open MoE model beating OLMo 3 and Apertus 70B on English and German benchmarks while activating only 3.2B of 31.6B parameters.
InternVLA-A1.5 Unifies Vision, Foresight, Action — SOTA on All Six Sim Benchmarks
InternVLA-A1.5 unifies vision-language understanding, latent foresight, and action into one robot policy, achieving SOTA on all six simulation benchmarks.
Claude Code Tops JetBrains' New Kotlin Benchmark with 85.7% Resolution
Claude Code with Opus 4.7 xhigh tops JetBrains' Kotlin Benchmark at 85.7%. Configure your CLAUDE.md with Kotlin conventions and use `--model opus-4.7-xhigh` to match this performance.
LLMForge: 7 Models Score 0.89 on CAD Benchmark; VLMs Fix Cylinders
LLMForge scores 0.89 on 97-design CAD benchmark. VLM critic achieves 100% watertight meshes but cylinders remain a failure mode.
Ant Group's 1.1B LingBot-Vision Beats Meta's 7B DINOv3 on 12 Benchmarks
Ant Group's 1.1B LingBot-Vision tops Meta's 7B DINOv3 on 12 spatial benchmarks, with 40% fewer FLOPs.
DARPA AIQ Program Shifts From Benchmarks to Measuring AI Capabilities
DARPA AIQ program, one year in, shifts from benchmarks to a science of AI capability, per program lead @patrickshafto.
Zhipu GLM-5.2 beats Anthropic's Mythos on bug-hunt benchmark
Zhipu AI's GLM-5.2 beat Anthropic's Claude Opus 4.8 on a cybersecurity bug-hunting benchmark, then matched it with extra instructions, marking another 'DeepSeek moment'.
GPT-5.6 Sol, Terra, Luna: Benchmark Performance Depends on Which Test You Use
OpenAI released GPT-5.6 as three tiers—Sol, Terra, Luna—on June 27, 2026. Sol tops Terminal-Bench 2.1 but trails competitors on other benchmarks. The release shifts focus to tiered pricing and efficiency, but access remains restricted.
SciCode: Epoch AI Launches Benchmark Measuring AI Research Ability
Epoch AI launched SciCode benchmark testing LLMs on real research coding tasks. Top models score below 30%, exposing gap between coding benchmarks and scientific ability.
Epoch AI's CursorBench Benchmarks AI Code Editing at Scale
Epoch AI launched CursorBench, a 500-task benchmark for AI code editors. It reveals a 15% accuracy gap vs. humans and 3x latency variance.
MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits
Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.
Zhipu GLM-5.2 tops global coding benchmarks, sparks 'DeepSeek moment'
Zhipu AI's GLM-5.2 ranks top-3 globally on a coding benchmark, with US engineers calling it a daily driver superior to GPT-5.5.
OpenAI GPT-5.5-Cyber Beats Anthropic Mythos on Security Benchmarks
OpenAI's GPT-5.5-Cyber beats Anthropic's Mythos on security benchmarks. Updated Codex plugin auto-patches after scanning 30M commits.
OpenAI shows small doses of beneficial-trait RL improve 44 of 53 safety benchmarks — and the gains generalize
OpenAI researchers Jagadeesh, Saab, Singhal et al. published findings on June 18 showing RL training on traits like honesty and corrigibility improved 44 of 53 safety benchmarks. Gains generalized across domains not used in training, and the model resisted harmful fine-tuning better than the baselin
Estonian Institute: Claude Tops Russian Propaganda Benchmark, Mistral Trails
Estonian Language Institute benchmark tests 60 AI models vs Russian propaganda. Claude tops, Mistral trails with 36.67% misinformation rate.
Visual-Seeker: Active Visual Reasoning Beats Proprietary MLLMs on 5 Benchmarks
Visual-Seeker achieves SOTA on five multimodal search benchmarks, surpassing proprietary models by actively harvesting visual evidence during search.
NVIDIA Blackwell Ultra Leads First Agentic AI Benchmark, 20x Agents/MW vs Hopper
NVIDIA Blackwell Ultra NVL72 leads the first AgentPerf benchmark for agentic AI, delivering 20x more agents per megawatt than Hopper.