On July 21, Google shipped Gemini 3.6 Flash. It scored 83.0% on OSWorld-Verified, beating GPT-5.6 Luna and grok-4-5" class="entity-chip">Grok 4.5.
Key facts
- 83.0% on OSWorld-Verified computer-use benchmark.
- $7.50 per million output tokens — cheapest among top scorers.
- Shipped July 21, 2026 by Google.
- Beats GPT-5.6 Luna and Grok 4.5.
- Flash tier traditionally trades capability for cost.
On July 21, Google shipped Gemini 3.6 Flash. On OSWorld-Verified — the benchmark that measures whether an agent can actually drive a real desktop, click the right buttons, fill the right fields, and finish a multi-step task — it scored 83.0%. That is ahead of GPT-5.6 Luna. Ahead of Grok 4.5. Ahead of Google’s own previous-generation flagship. And it did it as a Flash model: the cheap, fast tier that exists to be the thing you call ten thousand times a day when you can’t afford to call the expensive one.
I have spent two years watching the “cheap tier” lose every benchmark that mattered and win only on price. The whole point of a Flash or a Mini or a Haiku was that you traded capability for cost. Gemini 3.6 Flash broke that trade on the one benchmark where I’d have bet money the trade still held. Here’s what actually happened, the exact numbers, and runnable code so you can point it at your own desktop in about five minutes.
How the Cheap Tier Broke the Trade-Off

A model that costs $7.50 per million output tokens just posted the highest computer-use score in the industry. Not the highest cheap score. The highest score, period. This inverts the standard capability-cost curve. For reference, GPT-5.6 Luna costs roughly $30 per million output tokens [According to the source]. Grok 4.5 pricing is undisclosed but sits in the premium tier. The Flash model achieves 83.0% at less than a third of the cost of its closest competitor.
What OSWorld-Verified Actually Measures
OSWorld-Verified requires an agent to navigate a real desktop environment — clicking, typing, scrolling — across multi-step tasks like installing software, editing files, or filling out forms. It's not a synthetic QA benchmark; it's a proxy for real-world automation. The 83.0% score implies that for every 100 tasks, Gemini 3.6 Flash completes 83 without human intervention. That's production-grade reliability for enterprise automation workflows.
Implications for the Agent Market

Google's Gemini line now spans from the cheap Flash to the premium 3 Pro, with GPT-5 and Grok as direct competitors. If a Flash model can lead on computer use, the premium models must justify their price with something else — perhaps reasoning depth or multimodal understanding. The benchmark suggests that for task automation, cost is no longer a proxy for quality.
What to watch
Watch for Google's next earnings call (likely late October 2026) to see if enterprise adoption of Gemini 3.6 Flash for automation workflows accelerates, and whether OpenAI or xAI respond with price cuts on their premium models.
Source: pub.towardsai.net









