Google DeepMind launches Gemini Robotics ER 2 for complex embodied reasoning
Google DeepMind has introduced Gemini Robotics ER 2, an embodied reasoning model that leverages multimodal video understanding and latency-optimized APIs to orchestrate complex robotic tasks. The release, which integrates with existing Gemini developer platforms, is aimed at improving real-time spatial awareness and multi-robot collaboration for industrial and physical applications.
Key Takeaways
- Gemini Robotics ER 2 achieved 91.3% accuracy in precision moment-finding, determining the exact frame to transition between tasks.
- The model enables multi-robot collaboration, allowing diverse hardware like Apptronik’s Apollo 2 and Franka F3 Duo to share semantic task maps.
- Video-based success detection now monitors raw feeds for mid-execution failures such as spills or misalignments, replacing static snapshot verification.
- Integrated safety benchmarks autonomously halt humanoid platforms during human proximity and resume tasks only once the area is clear.
Why It Matters
The launch marks a transition from simple object recognition to temporal intelligence, where robots can reason through sequences lasting several minutes rather than isolated actions. By decoupling high-level reasoning from vision-language-action (VLA) motor control, Google allows developers to use lower-compute action models for movement while relying on Gemini for complex orchestration. This reduces the 'stop-and-think' latency that has historically hampered real-time physical AI. For the ecosystem, this establishes a standardized reasoning layer that could unite fragmented hardware platforms under a common semantic framework. Industry observers should watch for integration rates in third-party hubs like the Gemini Enterprise Agent Platform to gauge developer adoption.
Additional Context
The release of Gemini Robotics ER 2 follows a strategic expansion of Google DeepMind’s physical AI partnerships. Per The Robot Report, June 2026, Google significantly increased its footprint in specialized data collection by supporting Apptronik’s 'Robot Park' facility in Austin. This 90,000-square-foot site is used to generate the multimodal datasets necessary for training embodied models on delicate tasks, such as handling irregular objects and operating digital displays. These advancements helped address the 'dexterity gap' that previously limited humanoids to basic logistics.
Simultaneously, Google is tightening its technical collaboration with Boston Dynamics. Per B2B Tech News, July 2026, Boston Dynamics integrated Gemini Robotics models directly into its Orbit AIVI-Learning platform to reduce the man-hours required for manual teleoperation. This integration makes cloud connectivity a core operational dependency for industrial deployments of the Spot quadruped, reflecting a broader market shift where hardware performance is increasingly dictated by the latency of off-board foundation models.
Competition in the humanoid sector reached a new peak in mid-2026. While Google provides the intelligence layer, rivals like Figure have focused on vertical integration. According to RoboZaps, July 2026, Figure delivered over 350 of its Figure 03 units by early Q2, utilizing its proprietary Helix 02 whole-body AI stack. The market is now split between manufacturers building end-to-end proprietary systems and partners like Apptronik and Mercedes-Benz who are adopting Google’s Gemini as a universal high-level orchestrator for varied industrial fleets.
Read full article at blog.google
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source