Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Dashboard comparing AI model benchmark scores, with Locus bar exceeding Qwen3 bar, charts and metrics visible
AI ResearchScore: 87

Intology's Locus beats human-tuned Qwen3-1.7B in auto post-training

Intology's Locus beat human-tuned Qwen3-1.7B (51.6% vs 49.4%) on PostTrainBench by scaling compute 64x, showing AI research agents need longer timescales.

·11h ago·3 min read··15 views·AI-Generated·Report error
Share:
How did Intology's Locus beat the human-tuned Qwen3-1.7B on PostTrainBench?

Intology's automated research system Locus post-trained a model scoring 51.6% on PostTrainBench, beating the official human-tuned Qwen3-1.7B at 49.4%. Locus ran 100 hours across a cluster (4,500 H100-hours total) versus the benchmark's single H100 and 10-hour limit.

TL;DR

Locus scores 51.6% vs 49.4% on PostTrainBench · Removed H100-hour cap, raised budget 64x · Automated research beats human tuning at scale

Intology's Locus scored 51.6% on PostTrainBench, beating the official human-tuned Qwen3-1.7B's 49.4%. The win came from removing the benchmark's compute cap, scaling from 70 to 4,500 H100 hours per run.

Key facts

  • Locus scored 51.6% vs 49.4% for human-tuned Qwen3-1.7B
  • PostTrainBench+ budget: 4,500 H100-hours vs 70 baseline
  • Agent runs 100 hours across a cluster, not 10 on one H100
  • First public AI agent beats human post-training tuners
  • Single run per setting; stability not disclosed

Intology's automated research system, Locus, has post-trained a model that beats the official human-tuned Qwen3-1.7B release, scoring 51.6% against 49.4% According to @rohanpaul_ai. The result is the first public demonstration of an AI agent outperforming human researchers at the specific task of improving another model's post-training.

The benchmark and its constraints

PostTrainBench is the benchmark used to score AI agents that post-train other models, and it gives each agent one H100 and 10 hours. That is a tight budget — roughly 10 H100-hours per run — designed to measure how efficiently an agent can navigate the post-training pipeline.

Intology's different path was to remove that limit and let agents run for 100 hours across a cluster. So Intology raised the ceiling and called it PostTrainBench+, taking the total budget from 70 H100 hours to 4,500. The 64x increase in compute changes what the benchmark measures: not just efficiency, but the ability to sustain a long experimental loop.

The capability being measured here is sustained experimental search: allocating compute, running parallel jobs, reading evaluations, abandoning weak branches, and scaling the promising ones. At 10 hours, an agent can barely complete one or two training runs. At 100 hours across multiple GPUs, it can iterate dozens of times, learning from each evaluation.

What this means for automated research

The result suggests that the bottleneck for AI-driven research is not algorithmic cleverness but timescale. When given the compute budget to actually explore, Locus found configurations that human tuners missed. The 2.2 percentage point gap on PostTrainBench is modest in absolute terms, but it is the direction of the delta that matters — an AI system outperformed the human baseline on a task that was previously considered human expertise.

One run per setting is a serious limitation. The source does not disclose how many seeds were run, whether the result is stable across random initializations, or how Locus compares to other automated research systems under the same relaxed budget. The company did not disclose the full methodology.

Still, this makes a good case that automated research systems need to be evaluated at the timescale where research decisions compound. If the trend holds, the next generation of post-training pipelines may be designed by AI agents that run for weeks, not by humans working in days.

What to watch

Watch for Intology to publish replication runs across multiple seeds, and for PostTrainBench maintainers to respond — either by raising the official compute cap or adding a 'research timescale' track. If Locus's lead holds across seeds, expect competitors to follow with similar long-horizon research agents within two quarters.

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The result is less about Locus being clever and more about the benchmark's artificial constraint. PostTrainBench was designed to measure efficiency under scarcity — one H100, ten hours. That rewards agents that can guess well, not agents that can explore thoroughly. By removing the cap, Intology changed the measurement from 'how well can you optimize under a tight budget' to 'how well can you run a research lab.' Those are different skills, and the field has not yet decided which one matters more for real-world post-training. The 2.2 point delta is small enough that it could be noise, and the single-run methodology is a genuine weakness. But the direction is consistent with a broader pattern: as compute becomes cheaper relative to human researcher time, the optimal allocation shifts toward letting machines explore more. The comparison to prior art is stark — earlier automated ML systems like AutoML or NAS ran within fixed budgets and rarely beat human experts on end-to-end tasks. Locus's approach of removing the budget entirely is the structural break. The deeper question is whether this generalizes beyond Qwen3-1.7B. Beating a small open model's human-tuned release is one thing; beating a team of researchers working on a frontier model with weeks of compute is another. The source does not address this, and it is the gap that will determine whether this is a curiosity or a turning point.

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all