Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

A humanoid robot arm assembling electronic components on a lab bench, with a screen displaying video frames and…

Gemini Robotics ER 2 Hits 60% Video Completeness, Beats 1.6

Google's Gemini Robotics 2.0 ships ER 2 VLM with 60% video accuracy, but dexterity and safety models stay unreleased.

·11h ago·3 min read··10 views·AI-Generated·Report error
Share:
Source: arstechnica.comvia ars_technica_aiSingle Source
What did Google announce with Gemini Robotics 2.0?

Google's Gemini Robotics 2 includes three sub-models: ER 2, a dexterity model, and a safety model. Only Gemini Robotics ER 2 is publicly available via the Gemini Live API, claiming ~60% video frame completeness accuracy, up from the 1.6 release.

TL;DR

Google unveils Gemini Robotics 2, three sub-models. · ER 2 achieves ~60% video completeness accuracy. · Only ER 2 is public via Gemini Live API.

Google DeepMind unveiled Gemini Robotics 2 on July 30, 2026, with three sub-models, but only gemini-robotics-er-2" class="entity-chip">Gemini Robotics ER 2 is publicly available. The embodied reasoning VLM hits nearly 60 percent video frame completeness, up from the 1.6 release.

Key facts

  • Gemini Robotics 2.0: three sub-models, only ER 2 public
  • ER 2 achieves ~60% video frame completeness accuracy
  • ER 2 processes live video from robot cameras
  • Released via Gemini Live API, July 30, 2026
  • Dexterity and safety models not yet released

Google DeepMind's Gemini Robotics 2.0, announced July 30, 2026, brings a trio of sub-models aimed at generalist robots — machines that can handle any task a human could, what DeepMind scientists call "physical AGI." According to Ars Technica, the release includes an upgraded "embodied reasoning" model, a dexterity model, and a safety model. Only the first, Gemini Robotics ER 2, is available to developers now via the Gemini Live API.

The headline capability is live video processing. ER 2, a vision language model (VLM), can now track progress from the robot's cameras in real time, allowing the system to adjust as it moves from step to step. Google claims ER 2 classifies video frame completeness with almost 60 percent accuracy — far from perfect, but better than the 1.6 release or competing models' visual understanding. The company did not disclose benchmark details or training costs.

What's Actually New

The dexterity and safety models remain unreleased, with Google staying silent on timelines. That's a notable gap: the press release touts "improved dexterity" and "safety," yet developers can only test the reasoning layer. The two withheld models are where the claimed leaps in humanoid hand control and safe operation would land. Google's decision to gate them suggests either they're not production-ready or the company is holding back its best robotics assets.

Why This Matters

This release lands as Google's AI capex hits record levels — the company posted its first negative free cash flow since 2004 in late July, per our reporting. Robotics is a logical outlet for that spend, but the 60 percent video accuracy figure is a sobering counterpoint. It's a reminder that embodied AI remains in early innings; the gap between demo videos of backflipping robots and reliable generalist operation is still wide. For developers, ER 2 via the Gemini Live API is a concrete tool to test, but the real prize — the dexterity and safety models — stays locked behind Google's roadmap.

What to watch

Watch for Google's release timeline for the dexterity and safety sub-models. If they arrive within two quarters, expect a push into humanoid robot control. Also track developer adoption of ER 2 via the Gemini Live API and any benchmark comparisons against competitors like AgiBot's WITA-Omni, which scored 85.21 on DailyOmni.


Source: arstechnica.com


Sources cited in this article

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from 1 verified source, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

Google's decision to ship only the reasoning model is a strategic tell. The 60 percent accuracy figure, while improved, is a far cry from production-grade reliability. This suggests Google is using developers as beta testers for its embodied reasoning stack, while keeping the more differentiated dexterity and safety models close. The competitive landscape is heating up: AgiBot's WITA-Omni recently scored 85.21 on DailyOmni, beating Gemini in that benchmark. Google's robotics push is also a hedge against its capex spiral — the company posted negative free cash flow in Q2 2026, and robotics could eventually justify that spend, but only if the models deliver. The live video processing is the real differentiator here. Most VLMs operate on static images; ER 2's ability to process continuous video feeds is a meaningful architectural step. But the 60 percent accuracy is a reminder that we're still in the 'demo-ware' phase of embodied AI. The backflip videos are polished, but the underlying reliability is not there yet. Google's claim of 'physical AGI' remains aspirational; the gap between a 60% frame-completeness classifier and a robot that can 'do anything a human could do' is vast. The withheld models are the more interesting story. If the dexterity model can actually control complex humanoid hands, that's a step beyond what most competitors offer. But Google hasn't shown it, and until it does, the press release is just marketing. The company's integration with the Gemini Live API is smart — it lowers the barrier for developers to experiment, but it also means Google gets real-world feedback on its weakest model first.
Compare side-by-side
Gemini Robotics 2.0 vs Gemini Live API
Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all