Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Three small Google Gemini Flash model boxes labeled 1.5, 2.0, 2.5 sit on a shipping dock, with a large empty space…

Google Ships 3 Flash Models as 3.5 Pro Remains Missing

Google shipped three Gemini Flash models but 3.5 Pro remains delayed. Efficiency gains don't close the frontier gap with OpenAI and Anthropic.

·17h ago·4 min read··13 views·AI-Generated·Report error
Share:
Source: the-decoder.comvia the_decoderSingle Source
What new Gemini models did Google ship and why is 3.5 Pro still missing?

Google released three Gemini Flash models (3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber) but its flagship 3.5 Pro remains delayed. Gemini 4 pretraining is underway. The Flash models target efficiency and security, not frontier performance.

TL;DR

Gemini 3.6 Flash cuts token usage by 65%. · 3.5 Pro delayed; Gemini 4 pretraining underway. · Google trails OpenAI, Anthropic at frontier level.

Google shipped three Gemini Flash models on July 28, 2026, but its flagship 3.5 Pro remains delayed. The company is losing frontier ground to OpenAI, Anthropic, and Chinese labs while gemini-4" class="entity-chip">Gemini 4 pretraining begins.

Key facts

  • Gemini 3.6 Flash cuts token usage by 65% on DeepSWE.
  • Price: $1.50/M input, $7.50/M output tokens.
  • 3.5 Pro delayed; Gemini 4 pretraining underway.
  • 3.5 Flash-Lite: 350 output tokens/second.
  • DeepSWE improved from 37% to 49% over 3.5 Flash.

Google has announced three new models in the Gemini Flash family: 3.6 Flash, 3.5 Flash-Lite, and the cybersecurity model 3.5 Flash Cyber. Buried in the announcement is the fact that Google's anticipated flagship, Gemini 3.5 Pro, is still being tested exclusively with partners and will ship "as soon as it is ready." According to The Decoder

Google says pretraining for Gemini 4 is already underway, calling it its "most ambitious training run" yet. That reads like damage control, and it makes clear that Google knows what the market expects but can't deliver yet.

Key Takeaways

  • Google shipped three Gemini Flash models but 3.5 Pro remains delayed.
  • Efficiency gains don't close the frontier gap with OpenAI and Anthropic.

Efficiency gains, not frontier leaps

Gemini 3.6 Flash is expected to use about 17 percent fewer output tokens than 3.5 Flash, with savings reaching 65 percent on benchmarks like DeepSWE. Google has cut the price to $1.50 per million input tokens and $7.50 per million output tokens, making it much cheaper than the earlier 3.1 Pro model, which 3.6 Flash consistently beats in benchmarks.

Google also reports gains over 3.5 Flash: DeepSWE rises from 37 to 49 percent, MLE Bench from 49.7 to 63.9 percent, and OSWorld-Verified from 78.4 to 83 percent. The GDPval-AA v2 knowledge work benchmark improves from 1,349 to 1,421 points. Computer Use is now a built-in client-side tool in the Gemini API and Gemini Enterprise. Google has also added stronger Frontier Safety safeguards against CBRN misuse and cyberattacks.

Despite gains on multimodal tasks and a one million token context window, Google still trails the best models from competitors in the US and China. Logan Kilpatrick, a member of the technical staff, responded to criticism on X, saying the explicit goal was efficiency, usability, and lower cost, and that performance still improved in the process.

The missing frontier model

Gemini 3.5 Flash-Lite is tuned for low latency and high throughput, producing 350 output tokens per second at $0.30 per million input tokens. The 3.5 Flash Cyber model is restricted to governments and select partners due to potential risks. But the absence of a Pro-class model leaves Google without a competitive answer to OpenAI's GPT-4o or Anthropic's Claude Opus 4.6. [According to The Decoder], without the Pro model, Google is losing ground to OpenAI, Anthropic, and even Meta, which are pushing ahead with more capable frontier models.

Zwei Balkendiagramme zum durchschnittlichen Output-Token-Verbrauch pro Aufgabe: In DeepSWE v1.1 fällt er von 276K auf 97K, im Artificial Analysis Inte

Google's strategy appears to be focusing on efficiency and specialized use cases rather than chasing benchmark-topping performance. But that's a risky bet when enterprise customers increasingly demand frontier reasoning for complex workflows. The company's history shows it can catch up—Gemini 2.5 Pro was a breakthrough—but the current gap is widening.

What to watch

Watch for Google's next earnings call (likely October 2026) for any update on 3.5 Pro release timing or Gemini 4 pretraining milestones. Also track whether enterprise customers start migrating to OpenAI or Anthropic for reasoning-heavy workloads.

Vier Balkendiagramme vergleichen Gemini 3.1 Pro, 3.5 Flash und 3.6 Flash in DeepSWE v1.1, MLE-Bench, GDPVal-AA v2 und OSWorld-Verified; 3.6 Flash lieg


Source: the-decoder.com


Sources cited in this article

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

Google's strategy of shipping Flash models while 3.5 Pro remains missing is a calculated risk that reveals structural constraints. The company is effectively conceding the frontier benchmark race for now, betting that efficiency and specialized models (like the government-only Cyber variant) will retain enterprise customers. But this is a dangerous game: without a Pro-class model, Google can't compete on the reasoning-heavy workloads that drive high-value enterprise contracts. The fact that Gemini 4 pretraining is already underway suggests internal recognition that 3.5 Pro may never ship as a competitive product—it may be a bridge to something better. Compared to the current landscape, OpenAI and Anthropic are shipping frontier models at a faster cadence, with Anthropic's Claude Opus 4.6 and OpenAI's GPT-4o both available broadly. Google's Flash models are competitive on cost-per-token and specific benchmarks like DeepSWE, but they don't match the general reasoning capability of frontier models. The 3.5 Flash Cyber model is interesting as a government-specific play, but it's a niche that won't move the needle on Google's overall AI market share. The unique take here is that Google may be deliberately de-emphasizing the frontier race in favor of a 'good enough + cheap' strategy, mirroring what AWS did with Graviton chips. But in AI, where performance directly translates to capability for complex tasks, this could backfire if customers decide the gap is too wide.
Compare side-by-side
Anthropic vs Google
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all