Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

benchmark

30 articles about benchmark in AI news

Hugging Face Roundup: 10 Papers Push Harder Benchmarks, Multi-Agent AI

HF's weekly list of 10 papers signals a shift to harder benchmarks and multi-agent AI. Includes SWE-Bench ProMax and Alaya-EVOKE.

75% relevant

SPOT Distillation Beats OPD, EOPD on Reasoning Benchmarks

SPOT distillation beats OPD and EOPD on reasoning benchmarks by sparsely probing key positions and using outcome-calibrated targets. Reported by @HuggingPapers.

85% relevant

StreamArena: 243-Video Benchmark for Hour-Long Streaming AI

StreamArena launches with 243 hour-long videos and 3,646 open-ended tasks for interactive streaming video understanding, pushing beyond short-clip benchmarks. No baselines yet.

85% relevant

Agent Memory Benchmarks Get a Shared 5,000-Question Test

A 20+ institution consortium launched a shared agent memory benchmark with 5,000 questions and one answer model, targeting the attribution problem in memory startup claims. First rankings due mid-August.

78% relevant

Supabase's Evals Benchmark Just Gave Claude Code a Real-World Report Card

Supabase Evals is an open-source benchmark that scores Claude Code, Codex, and OpenCode on real Supabase tasks. Run `supabase eval` on your repo to find agent weaknesses.

100% relevant

ClBench-V: New Benchmark Tests Multimodal Contextual Learning in 3 Dimensions

ClBench-V benchmark from @HuggingPapers tests multimodal contextual learning across three dimensions: grounding, application, and knowledge learning. No results disclosed yet.

85% relevant

Ant Ling-3.0-flash Beats 1T-Ring-2.6 on 11 of 12 Benchmarks

Ant's 124B-param Ling-3.0-flash with 5.1B activated beats 1T-Ring-2.6 in 11 of 12 benchmarks, tying DeepSeek V4 Flash. Sparse activation economics are the story.

90% relevant

FutureX Refactoring Benchmark: 40% Faster Than Claude Code, 80% Test Pass Rate

FutureX refactored code 40% faster than Claude Code in a controlled benchmark, with an 80% initial test pass rate vs 60%. The specialized agent required 4 minutes of review per task versus 7 minutes for Claude Code.

95% relevant

Benchmark lets image models answer in pixels, not text

New 'Show, Don't Tell' benchmark tests spatial cognition via pixel-level outputs. GPT Image 2 solves 37% of cases missed by GPT-5.4, highlighting a gap in text-based spatial reasoning.

85% relevant

ActiveVision Benchmark: Humans 96.1%, Best AI 10.6%

ActiveVision benchmark: humans 96.1%, best AI 10.6%. The 85.5-point gap reveals fundamental limits in iterative visual reasoning for current models.

85% relevant

90 Hours of Black Myth: Wukong Fuel New World Model Benchmark

A new survey and benchmark rethinks interactive world models as game engines, with a data engine collecting over 90 hours of Black Myth: Wukong gameplay.

78% relevant

KeyFrame-Compass Benchmark Targets Keyframe Video Generation Gaps

KeyFrame-Compass is the first benchmark for keyframe-conditioned video generation, with 386 samples and six metrics.

80% relevant

gdb: Benchmarks Saturate Too Fast for Reliable AI Progress Tracking

@gdb notes benchmarks saturate quickly. This undermines AI progress tracking and may force shift to dynamic evaluations.

75% relevant

Soofi S 30B-A3B: German open model tops English, German benchmarks

German consortium releases Soofi S 30B-A3B, an open MoE model beating OLMo 3 and Apertus 70B on English and German benchmarks while activating only 3.2B of 31.6B parameters.

100% relevant

InternVLA-A1.5 Unifies Vision, Foresight, Action — SOTA on All Six Sim Benchmarks

InternVLA-A1.5 unifies vision-language understanding, latent foresight, and action into one robot policy, achieving SOTA on all six simulation benchmarks.

85% relevant

Claude Code Tops JetBrains' New Kotlin Benchmark with 85.7% Resolution

Claude Code with Opus 4.7 xhigh tops JetBrains' Kotlin Benchmark at 85.7%. Configure your CLAUDE.md with Kotlin conventions and use `--model opus-4.7-xhigh` to match this performance.

98% relevant

LLMForge: 7 Models Score 0.89 on CAD Benchmark; VLMs Fix Cylinders

LLMForge scores 0.89 on 97-design CAD benchmark. VLM critic achieves 100% watertight meshes but cylinders remain a failure mode.

75% relevant

Ant Group's 1.1B LingBot-Vision Beats Meta's 7B DINOv3 on 12 Benchmarks

Ant Group's 1.1B LingBot-Vision tops Meta's 7B DINOv3 on 12 spatial benchmarks, with 40% fewer FLOPs.

100% relevant

DARPA AIQ Program Shifts From Benchmarks to Measuring AI Capabilities

DARPA AIQ program, one year in, shifts from benchmarks to a science of AI capability, per program lead @patrickshafto.

75% relevant

Zhipu GLM-5.2 beats Anthropic's Mythos on bug-hunt benchmark

Zhipu AI's GLM-5.2 beat Anthropic's Claude Opus 4.8 on a cybersecurity bug-hunting benchmark, then matched it with extra instructions, marking another 'DeepSeek moment'.

93% relevant

GPT-5.6 Sol, Terra, Luna: Benchmark Performance Depends on Which Test You Use

OpenAI released GPT-5.6 as three tiers—Sol, Terra, Luna—on June 27, 2026. Sol tops Terminal-Bench 2.1 but trails competitors on other benchmarks. The release shifts focus to tiered pricing and efficiency, but access remains restricted.

76% relevant

SciCode: Epoch AI Launches Benchmark Measuring AI Research Ability

Epoch AI launched SciCode benchmark testing LLMs on real research coding tasks. Top models score below 30%, exposing gap between coding benchmarks and scientific ability.

95% relevant

Epoch AI's CursorBench Benchmarks AI Code Editing at Scale

Epoch AI launched CursorBench, a 500-task benchmark for AI code editors. It reveals a 15% accuracy gap vs. humans and 3x latency variance.

95% relevant

MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits

Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.

77% relevant

Zhipu GLM-5.2 tops global coding benchmarks, sparks 'DeepSeek moment'

Zhipu AI's GLM-5.2 ranks top-3 globally on a coding benchmark, with US engineers calling it a daily driver superior to GPT-5.5.

100% relevant

OpenAI GPT-5.5-Cyber Beats Anthropic Mythos on Security Benchmarks

OpenAI's GPT-5.5-Cyber beats Anthropic's Mythos on security benchmarks. Updated Codex plugin auto-patches after scanning 30M commits.

100% relevant

OpenAI shows small doses of beneficial-trait RL improve 44 of 53 safety benchmarks — and the gains generalize

OpenAI researchers Jagadeesh, Saab, Singhal et al. published findings on June 18 showing RL training on traits like honesty and corrigibility improved 44 of 53 safety benchmarks. Gains generalized across domains not used in training, and the model resisted harmful fine-tuning better than the baselin

95% relevant

Estonian Institute: Claude Tops Russian Propaganda Benchmark, Mistral Trails

Estonian Language Institute benchmark tests 60 AI models vs Russian propaganda. Claude tops, Mistral trails with 36.67% misinformation rate.

72% relevant

Visual-Seeker: Active Visual Reasoning Beats Proprietary MLLMs on 5 Benchmarks

Visual-Seeker achieves SOTA on five multimodal search benchmarks, surpassing proprietary models by actively harvesting visual evidence during search.

72% relevant

NVIDIA Blackwell Ultra Leads First Agentic AI Benchmark, 20x Agents/MW vs Hopper

NVIDIA Blackwell Ultra NVL72 leads the first AgentPerf benchmark for agentic AI, delivering 20x more agents per megawatt than Hopper.

92% relevant