Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Dashboard showing InAgent benchmark score 90.2% on OSWorld, with comparison bars for OpenAI, Google, and Anthropic
AI ResearchBreakthroughScore: 90

InAgent Hits 90.2% on OSWorld, First Agent Past 90%

InAgent scored 90.2% on OSWorld, first above 90%, with 100% on system-level tasks, surpassing OpenAI, Google, and Anthropic records. Harness engineering, not raw model power, drove the result.

·10h ago·3 min read··21 views·AI-Generated·Report error
Share:
Source: pandaily.comvia pandailyCorroborated
What is InAgent's OSWorld score and how does it compare to OpenAI, Google, and Anthropic?

InAgent, a Chinese computer-use agent, scored 90.2% task success on OSWorld in July 2026, the first agent above 90%, with 100% on system-level tasks. The result surpasses records from OpenAI, Google, and Anthropic, highlighting harness engineering as the new competitive frontier.

TL;DR

InAgent scores 90.2% on OSWorld, first above 90% · 100% success on system-level tasks · Harness engineering becomes new AI battleground

InAgent scored 90.2% on OSWorld in July 2026, the first computer-use agent above 90%. The Chinese agent's result surpasses public records from OpenAI, Google, and Anthropic on the same benchmark.

Key facts

  • 90.2% OSWorld task success, first above 90%
  • 100% success on system-level tasks
  • Surpasses OpenAI, Google, Anthropic records
  • Single run; seed variance not disclosed
  • Harness engineering drives the result

InAgent scored 90.2% task success on OSWorld in July 2026, the first computer-use agent above 90%, according to Pandaily. The agent achieved 100% on system-level tasks, a category that has historically dragged down frontier models. OSWorld measures an agent's ability to complete real-world computer tasks — file operations, web browsing, application control — through screenshots and keyboard/mouse actions.

Key Takeaways

  • InAgent scored 90.2% on OSWorld, first above 90%, with 100% on system-level tasks, surpassing OpenAI, Google, and Anthropic records.
  • Harness engineering, not raw model power, drove the result.

Why the harness matters

The headline number obscures the structural story: the gap between frontier models has narrowed to the point where the scaffold around the model now determines benchmark placement. Harness engineering — the code that plans, verifies, and recovers from errors — has become the new AI competition frontier. InAgent's 90.2% is not primarily a model win; it is a systems win.

This mirrors what Supabase's evals benchmark showed in August 2026: real-world agent performance tracks tooling and orchestration more than raw model capability. Claude Code's strong showing there came from its scaffolding, not just Claude's weights. InAgent confirms the pattern on a harder, more standardized benchmark.

The numbers behind the record

What Screen Agent’s #1 OSWorld ranking means for UI automation i…

The 90.2% figure represents a single run. The source does not disclose variance across seeds, a meaningful omission for a benchmark where stochastic sampling can swing results by several points. OSWorld's system-level tasks — which require multi-step operations like installing software or configuring settings — are where most agents fail; InAgent's 100% there is the more impressive number.

Prior public records on OSWorld sat below 90%. OpenAI, Google, and Anthropic have each published computer-use agents, but none crossed the threshold. InAgent's result places the Chinese lab ahead on this specific metric, though the benchmark is narrow: OSWorld covers desktop tasks, not the full range of enterprise workflows where those companies compete.

What to watch

Watch whether InAgent publishes seed-variance data or a follow-up paper detailing its harness architecture. Also track whether OpenAI, Google, or Anthropic respond with OSWorld scores above 90% in the next two quarters — a response would confirm harness engineering as the new arms race.


Source: pandaily.com


Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

The 90.2% OSWorld score is less a model breakthrough and more a confirmation that agent performance now hinges on orchestration. Frontier labs have converged on model quality; the differentiator is the code that plans actions, verifies outcomes, and recovers from failures. InAgent's 100% on system-level tasks — where multi-step operations punish weak scaffolding — underscores this. This aligns with the Supabase evals benchmark from August 2026, which showed Claude Code's real-world performance tracking its tooling. The pattern is consistent: agent leaders are systems companies, not model companies. The source's silence on seed variance is the weak point — a single 90.2% run without error bars is a headline, not a scientific result. The competitive implication is direct. OpenAI, Google, and Anthropic have each published computer-use agents but none crossed 90%. If InAgent's harness is replicable and open-sourced, the moat shifts from weights to infrastructure. If it is closed, expect a scramble to reverse-engineer the approach.
Compare side-by-side
Anthropic vs OpenAI

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in AI Research

View all