Model Releases Google DeepMind Blog

Gemini Robotics 2 brings whole body intelligence to robots

Gemini Robotics 2VLADeepMindembodied AI

Most robots today are pre-programmed or teleoperated for narrow, repetitive task sequences, lacking the ability to learn or adapt to unpredictable environments, and transferring skills between robot bodies remains difficult. Gemini Robotics 2 is designed as an intelligence layer for the next generation of adaptable robots, building on Gemini's multimodal understanding to enable thinking, acting, and safe interaction in the real world.

The release includes three models: Gemini Robotics 2, a vision-language-action (VLA) model that converts vision and language input into motor control, capable of controlling full humanoids from feet to fingertips and other bi-arm robots, with advanced dexterous manipulation on both hands and grippers; Gemini Robotics ER 2, an embodied reasoning (ER) model that acts as an agent for communicating with humans, understanding the physical world, planning multi-step tasks lasting several minutes, and now supporting multi-robot teamwork; and Gemini Robotics On-Device 2, an efficient VLA optimized to run locally on robotic devices, achieving fast adaptation to completely new robot embodiments with a few hours of data.

These capabilities enable robots to reason through every movement, such as a humanoid walking, crouching, stretching, and manipulating objects to clean a cluttered room, even teaming up with other robots to finish tasks faster. The models can also seamlessly adapt to entirely new robotic bodies in just a few hours.

Gemini Robotics ER 2 is now available on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, while the VLA and On-Device models are available to early-access partners. Developers can find hardware integration guidance on the DeepMind Developer blog.

Read original →

← Back to home