Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Close-up of a sleek dark laptop screen displaying a terminal with code and benchmark metrics, surrounded by subtle…
AI ResearchScore: 85

StateM Hits 95.3% on Terminal-Bench 2.1 via Harness Scaling

StateM hits 95.3% on Terminal-Bench 2.1 via harness scaling, per Hugging Face roundup. Highlights evaluation infrastructure's impact on agent scores.

·1d ago·3 min read··48 views·AI-Generated·Report error
Share:
What is StateM's accuracy on Terminal-Bench 2.1 and how was it achieved?

StateM achieved 95.3% accuracy on Terminal-Bench 2.1 through harness scaling, a method that improves agent evaluation infrastructure rather than model weights. The result, highlighted in Hugging Face's weekly paper roundup, underscores how evaluation harness design can significantly inflate or improve agent performance scores.

TL;DR

StateM hits 95.3% on Terminal-Bench 2.1 · Harness scaling boosts agent accuracy · New agent benchmarks target video, embodied · SemaPLC grounds PLC code with verification

StateM hits 95.3% on Terminal-Bench 2.1 via harness scaling, per Hugging Face's weekly roundup. The result highlights how evaluation infrastructure—not model weights—can drive agent performance.

Key facts

  • StateM: 95.3% accuracy on Terminal-Bench 2.1
  • Method: harness scaling, not model weights
  • RA-Bench: AI-generated video detection in crises
  • SemaPLC: PLC code generation with verification gates
  • 10+ papers listed in weekly Hugging Face roundup

StateM's 95.3% accuracy on Terminal-Bench 2.1, reported in Hugging Face's weekly paper roundup, is notable not for the model itself but for the method: harness scaling According to @HuggingPapers. This technique improves the evaluation harness—the scaffolding that connects an agent to a terminal environment—rather than the underlying policy. The implication: agent benchmarks often measure harness engineering as much as agent intelligence.

Harness scaling is a known lever. Prior work on SWE-Bench showed that better environment setup, retry logic, and tool definitions can shift scores by 10-20 points without changing the model. StateM's result, if replicated, suggests Terminal-Bench 2.1 is similarly sensitive to harness design. The roundup does not disclose StateM's architecture, training compute, or the specific harness changes, leaving the result difficult to contextualize [Per the tweet].

Key Takeaways

  • StateM hits 95.3% on Terminal-Bench 2.1 via harness scaling, per Hugging Face roundup.
  • Highlights evaluation infrastructure's impact on agent scores.

Benchmark Proliferation

Paper page - StateM: Reaching 95.3% Raw Accuracy, or a $15 ...

Beyond StateM, the roundup lists several new benchmarks targeting agent and video domains. RA-Bench evaluates AI-generated video detection in real-world crises, addressing deepfake verification under chaotic conditions. SemComp-Bench measures semantic task completion in video generation, moving beyond pixel-level metrics. S^2VOPD boosts visual reasoning with self-supervised distillation, while EnvHarness reimagines static environments for agent learning—directly relevant to the harness-scaling theme.

These benchmarks share a pattern: they aim to fix evaluation blind spots. Terminal-Bench 2.1 itself was designed to test terminal command execution, a core agent skill. The proliferation suggests the field is converging on evaluation as a bottleneck for agent progress.

Verification and Embodied Learning

SemaPLC grounds PLC code generation with verification gates, ensuring generated code passes formal checks before deployment. This is a practical step for industrial automation, where unverified code is a safety risk. Zetta enables closed-loop embodied learning for physical intelligence, while HarnessEval-W brings agentified evaluation to world models—another harness-centric approach.

The trend is clear: harness engineering and verification are becoming first-class research topics. As agents move into production, the gap between model capability and reliable deployment narrows through better evaluation and control.

No training details, model sizes, or compute figures were provided for any of these papers in the roundup. The field should treat these results as preliminary until full papers and code are available.

What to watch

Watch for StateM's paper and code release—if the harness changes are open-sourced, expect replication efforts across other agent benchmarks. Also track whether Terminal-Bench 2.1 updates its leaderboard to control for harness variability, and whether RA-Bench's crisis video detection becomes a standard for deepfake verification.

Sources cited in this article

  1. Hugging Face's
  2. Hugging Face
Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 2 verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The StateM result is a textbook case of evaluation sensitivity. Prior work on SWE-Bench and similar agent benchmarks has shown that harness engineering can dominate model choice. If 95.3% is achievable via harness scaling alone, then Terminal-Bench 2.1's leaderboard is less a measure of agent capability and more a measure of who invested in evaluation tooling. This is not a criticism—harness engineering is legitimate research—but it demands scrutiny of any agent benchmark claim. The roundup's other entries reinforce this structural read. EnvHarness and HarnessEval-W explicitly target harness design, while SemaPLC's verification gates are a form of output-level control. The field is moving toward evaluation as a discipline, which is healthy but creates a new risk: benchmark gaming via harness overfitting. Researchers should demand ablation studies that isolate harness contributions, as StateM's tweet does not provide them. Contrarian take: The emphasis on harness scaling may be a sign that model-level agent improvements are plateauing. If the easiest path to higher scores is better scaffolding, then the frontier of agent research is shifting from architecture to infrastructure—a trend that favors companies with strong engineering teams over those with pure research chops.
This story is part of
Hugging Face Becomes the Neutral Ground Where Google and Anthropic's Agent Protocol War Converges
As Claude Code's MCP dominance threatens Google Cloud, Hugging Face's unique position as partner to both players creates an unexpected convergence zone

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all