Google shipped three Gemini Flash models on July 28, 2026, but its flagship 3.5 Pro remains delayed. The company is losing frontier ground to OpenAI, Anthropic, and Chinese labs while gemini-4" class="entity-chip">Gemini 4 pretraining begins.
Key facts
- Gemini 3.6 Flash cuts token usage by 65% on DeepSWE.
- Price: $1.50/M input, $7.50/M output tokens.
- 3.5 Pro delayed; Gemini 4 pretraining underway.
- 3.5 Flash-Lite: 350 output tokens/second.
- DeepSWE improved from 37% to 49% over 3.5 Flash.
Google has announced three new models in the Gemini Flash family: 3.6 Flash, 3.5 Flash-Lite, and the cybersecurity model 3.5 Flash Cyber. Buried in the announcement is the fact that Google's anticipated flagship, Gemini 3.5 Pro, is still being tested exclusively with partners and will ship "as soon as it is ready." According to The Decoder
Google says pretraining for Gemini 4 is already underway, calling it its "most ambitious training run" yet. That reads like damage control, and it makes clear that Google knows what the market expects but can't deliver yet.
Key Takeaways
- Google shipped three Gemini Flash models but 3.5 Pro remains delayed.
- Efficiency gains don't close the frontier gap with OpenAI and Anthropic.
Efficiency gains, not frontier leaps
Gemini 3.6 Flash is expected to use about 17 percent fewer output tokens than 3.5 Flash, with savings reaching 65 percent on benchmarks like DeepSWE. Google has cut the price to $1.50 per million input tokens and $7.50 per million output tokens, making it much cheaper than the earlier 3.1 Pro model, which 3.6 Flash consistently beats in benchmarks.
Google also reports gains over 3.5 Flash: DeepSWE rises from 37 to 49 percent, MLE Bench from 49.7 to 63.9 percent, and OSWorld-Verified from 78.4 to 83 percent. The GDPval-AA v2 knowledge work benchmark improves from 1,349 to 1,421 points. Computer Use is now a built-in client-side tool in the Gemini API and Gemini Enterprise. Google has also added stronger Frontier Safety safeguards against CBRN misuse and cyberattacks.
Despite gains on multimodal tasks and a one million token context window, Google still trails the best models from competitors in the US and China. Logan Kilpatrick, a member of the technical staff, responded to criticism on X, saying the explicit goal was efficiency, usability, and lower cost, and that performance still improved in the process.
The missing frontier model
Gemini 3.5 Flash-Lite is tuned for low latency and high throughput, producing 350 output tokens per second at $0.30 per million input tokens. The 3.5 Flash Cyber model is restricted to governments and select partners due to potential risks. But the absence of a Pro-class model leaves Google without a competitive answer to OpenAI's GPT-4o or Anthropic's Claude Opus 4.6. [According to The Decoder], without the Pro model, Google is losing ground to OpenAI, Anthropic, and even Meta, which are pushing ahead with more capable frontier models.

Google's strategy appears to be focusing on efficiency and specialized use cases rather than chasing benchmark-topping performance. But that's a risky bet when enterprise customers increasingly demand frontier reasoning for complex workflows. The company's history shows it can catch up—Gemini 2.5 Pro was a breakthrough—but the current gap is widening.
What to watch
Watch for Google's next earnings call (likely October 2026) for any update on 3.5 Pro release timing or Gemini 4 pretraining milestones. Also track whether enterprise customers start migrating to OpenAI or Anthropic for reasoning-heavy workloads.

Source: the-decoder.com








