Model Releases Google DeepMind Blog

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Gemini Robotics ER 2embodied reasoningroboticsmulti-robot collaboration

For robots to assist humans in everyday environments, accurate spatial reasoning is not enough—they must think fast and time their decisions with the real-time speed of the physical world. Google DeepMind introduces Gemini Robotics ER 2, its most capable 'embodied reasoning' model, acting as a high-level brain that lets robots chat, understand the physical world, plan multi-step tasks, and then hand off motor execution to any lower-level vision-language-action (VLA) model. It can also natively call tools like Google Search or any user-defined function.

The design allows the robot to 'think' about what comes next while simultaneously performing actions, using continuous video feeds to track its own progress, adapt if something goes wrong, and know when to move on. Multi-robot collaboration is introduced, enabling robots to work together in shared spaces and complete complex workflows a single robot could not handle alone. Developers can declare low-level control interfaces (VLA models, navigation APIs) as tools and stream video, audio, or text directly into the model, which integrates into the Gemini Live API via a bidirectional streaming endpoint optimized for latency-sensitive tasks.

Gemini Robotics ER 2 is a significant upgrade over ER 1.6, consistently outperforming it for tool orchestration across three control modes: real VLA, sim VLA, and human tele-op. The model is now publicly available via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform, with configuration examples provided to help developers power physical AI tasks.

Read original →

← Back to home