METR's new 'expenditure horizon' metric puts a dollar figure on when AI agents become cost-competitive with humans. On the NanoGPT speedrun, humans cost $2,500 per 1% speedup; AI agents range from $0 to $3,300.
Key facts
- Human cost: $2,500 per 1% speedup on NanoGPT.
- AI agent horizons: $0–$3,300 per run.
- METR interviewed two top NanoGPT contributors.
- Training time dropped from 45 min to under 2 min since May 2024.
- GPT-5.6 Sol scored 72.7% on DeepSWE on 2026-07-25.
Research organization METR proposes a new metric called the 'expenditure horizon' to answer a fundamental question: when does it become cheaper to pay a human than to run an AI agent? The metric compares the total cost—compute, API fees, and human labor—required to achieve the same improvement on a task. The expenditure horizon is the break-even point: below that budget, the AI is cheaper; above it, humans win According to The Decoder.
METR chose the NanoGPT speedrun as its testing ground—a community project where volunteers compete to train a language model as fast as possible on standardized hardware. Since May 2024, the required training time dropped from about 45 minutes to under two minutes across 82 documented improvement steps. To estimate human labor cost, METR interviewed two top contributors and used AI model Opus-4.6 to estimate effort per improvement. Both methods converged on roughly 16 hours of work per 1% speedup. At $150/hour, that's $2,500 per percentage point. METR stresses that this number is highly uncertain, noting that most time went into ideas that ultimately didn't work.
For the AI side, METR had six models work on the same task starting from an already optimized state (Record #78 from March 2026), with a budget cap of $10,000 per run. The results showed expenditure horizons between $0 and $3,300. The differences between models were stark: GPT-5 and Opus-4.1 pro outperformed others. Notably, GPT-5.6 Sol—which scored 72.7% on DeepSWE benchmark on 2026-07-25—was not explicitly tested in this study, but its capabilities suggest it could push the horizon higher. The metric has blind spots: it doesn't account for model improvements over time, and the task selection (NanoGPT speedrun) may not generalize to other domains. METR acknowledges these limitations, and the newest generation of models could change the picture.
Why the metric matters
Compared to typical AI benchmarks, the expenditure horizon has two advantages. First, it produces a fine-grained value showing how much improvement you get for how much money, rather than a pass-or-fail verdict. Second, it converts all costs into a single currency, covering compute, API costs, and human labor. This makes it directly applicable to real-world deployment decisions: a company can compute the exact budget at which it's cheaper to use an AI agent versus hiring a human contractor.
blind spots and model evolution
METR's metric is a step forward for cost-aware evaluation, but it has a fundamental blind spot: it compares current AI models against human baselines that are themselves moving targets. The $2,500 per 1% speedup for humans on NanoGPT comes from interviews with two contributors, but as the task becomes more optimized, the marginal human effort may increase or decrease. Moreover, the metric doesn't account for the value of AI agents that can run 24/7 without fatigue or the potential for models to improve via scaffold updates—a topic explored in a recent 239-paper survey According to gentic.news. The true expenditure horizon for production systems—where AI agents handle thousands of tasks simultaneously—could be far lower than these single-task experiments suggest.
Key facts
- Human cost: $2,500 per 1% speedup on NanoGPT.
- AI agent horizons: $0–$3,300 per run.
- METR interviewed two top NanoGPT contributors.
- Training time dropped from 45 minutes to under 2 minutes since May 2024.
- GPT-5.6 Sol scored 72.7% on DeepSWE benchmark on 2026-07-25.
What to watch
Watch for METR to release expenditure horizon results on more complex, multi-step software engineering tasks—especially those involving GPT-5.6 Sol and Claude Fable 5. If the horizon crosses $10,000 for a single task, it would signal that AI agents are becoming economically viable for high-value research and development work. Also watch for the Q3 ARR disclosures from Anthropic and OpenAI to see if enterprise seat counts cross 100K, which would indicate real-world adoption at scale.



Source: the-decoder.com
Key Takeaways
- METR's expenditure horizon metric shows AI agents break even at $0–$3,300 on NanoGPT, vs $2,500 per 1% speedup for humans.
- GPT-5 and Opus-4.1 pro lead, but blind spots remain.








