Google DeepMind Launches Gemini Robotics ER 2: The Safest, Most Capable AI Brain for Humanoid Control & Real-World Robots

🚀 Key Takeaways

  • Google DeepMind launched Gemini Robotics ER 2, its most capable embodied reasoning model for robotics.
  • Designed as a high-level brain, it represents a significant upgrade from its predecessor, ER 1.6.
  • The model enables intelligent whole-body control, real-time spatial reasoning, and multi-step task planning for robots.
  • Robots powered by Gemini Robotics ER 2 can self-correct, adapt to situations, and engage in multi-robot collaboration.
  • It significantly enhances safety features, making it Google's safest robotics model to date with improved human proximity detection.
  • This suite of AI models supports advanced capabilities like video understanding, human-robot chat, and native tool calling.
  • Developers can access Gemini Robotics ER 2 publicly via the Gemini API and Google AI Studio.
The landscape of artificial intelligence in robotics has taken a substantial leap forward with the recent announcement from Google DeepMind.
On July 30, 2026, the tech giant unveiled Gemini Robotics ER 2, marking a pivotal moment in the development of intelligent machines.
This latest iteration is hailed as Google's most capable embodied reasoning model designed specifically for robotics, promising to redefine how robots interact with and understand their physical environments.

As a comprehensive suite of AI models, Gemini Robotics ER 2 acts as a sophisticated high-level brain for a new generation of robots.
It significantly upgrades over its predecessor, Gemini Robotics ER 1.6, by integrating advanced multi-modal understanding and real-time reasoning capabilities.
This allows robots to perform complex tasks with unprecedented dexterity and intelligence, moving closer to seamless human-robot collaboration.

The model’s core strength lies in enabling robots to navigate, plan, and execute tasks with greater autonomy, including features like whole-body control and spatial reasoning.
Crucially, Gemini Robotics ER 2 also sets new benchmarks in safety, incorporating robust mechanisms to ensure robots can operate more securely in real-world scenarios, including enhanced human proximity awareness.
This launch not only accelerates the potential for helpful robots but also empowers developers with powerful tools to innovate further in the robotics domain.


1. Gemini Robotics ER 2: Google DeepMind's New Brain for Robots

This section provides a focused analysis of Gemini Robotics ER 2, the core reasoning model that powers the newly announced Gemini Robotics 2 system, connecting its specific advancements to the broader capabilities of the full-stack humanoid control platform.

An Overview of the Latest Release

Google DeepMind officially launched its latest robotics model, Gemini Robotics ER 2, on July 30, 2026.
The company has positioned it as the "most capable embodied reasoning model for robotics" currently available, signaling a major step forward in creating autonomous systems.

Key Features and Architectural Foundation

At its core, Gemini Robotics ER 2 is designed to function as a high-level brain for robots.
This new family of AI models, purposefully designed for robotics, is built upon the powerful foundation of Google's Gemini 2.0 architecture, which was first detailed in a 2025 report.
The system is not a single monolithic model but rather a suite of AI models that combines several different specialized components into one cohesive system, enabling more complex and nuanced task execution.

The Gemini Robotics 2 Family

The complete Gemini Robotics 2 family actually includes three distinct models, each likely tailored for different applications or scales.
However, as of the launch, Google DeepMind has only made one of these models publicly available for researchers and developers.
The specifics of the other two models in the suite have not yet been disclosed.


2. Unlocking Advanced Robotics: Core Capabilities of Gemini Robotics ER 2

This section delves into the foundational technologies of Gemini Robotics ER 2, the core AI model powering the advanced humanoid control systems Google announced on July 31st.

Enhanced Perception and Reasoning

Gemini Robotics ER 2 grants robots a sophisticated understanding of their environment through several key capabilities.
It fundamentally powers them with video understanding, allowing them to process and interpret visual information dynamically.
This is paired with real-time spatial reasoning, enabling a robot to comprehend its position and the arrangement of objects in the physical world around it.
Based on this perception, the system facilitates complex, multi-step task planning, breaking down a large goal into a sequence of achievable actions.
A significant leap forward is its support for multi-robot collaboration, which allows different types of machines to communicate and work together on a shared objective.
By processing continuous video feeds of their own actions, robots can even track their progress, creating a feedback loop for greater awareness.
This high-level reasoning allows the system to generalize its knowledge to more novel situations and reason through every movement, unlocking a much broader range of potential tasks.

Intelligent Task Execution and Self-Correction

The model's intelligence extends beyond planning to the actual execution and adaptation of tasks.
Gemini Robotics ER 2 orchestrates the necessary steps for a robot and critically enables self-correction.
This means robots can adapt if something goes wrong, effectively fixing their own mistakes in real time without human intervention.
The system's design is advanced enough to allow a robot to 'think' about what comes next while it is simultaneously performing a physical action.
This ensures a fluid and efficient workflow, as the robot knows exactly when to move on to the next step of its task.
Architecturally, Gemini Robotics ER 2 acts as a high-level "brain" that hands off the specific motor execution commands to lower-level vision-language-action (VLA) models.
It can also natively call external tools, such as performing a Google Search for information or accessing other user-defined functions to augment its capabilities.

Human-Robot Interaction and Whole-Body Control

A primary focus of Gemini Robotics ER 2 is enabling more natural and capable physical robots, particularly humanoids.
The system allows robots to chat with humans, providing a natural interface for receiving instructions and reporting status.
Crucially, it unlocks intelligent whole-body control, which is essential for complex machines.
For humanoids, this translates into the ability to walk and move naturally, navigating spaces with human-like fluidity.
Beyond basic locomotion, the model also enables advanced dexterity for manipulating objects.
This fine motor control allows humanoids to handle delicate tasks that were previously challenging, such as successfully tying knots.


3. Performance and Breakthroughs: How Gemini Robotics ER 2 Excels

This section delves into the specific performance metrics and technological advancements that define Google's July 31st announcement of Gemini Robotics ER 2, showcasing how the new model sets new benchmarks in robotic AI.

Benchmark Achievements and Efficiency

Gemini Robotics ER 2 demonstrates significant quantitative leaps in performance, establishing new standards for robotic task understanding and execution speed.
On progress classification tasks, the model achieves 57.4% accuracy, a figure that surpasses both previous generation models and competing frontier AIs.
Even more impressive are the results in moment-finding, a critical capability for identifying key steps in a process.
Here, ER 2 reaches 91.3% accuracy with a mean absolute distance of just 0.96 seconds, meaning it can pinpoint the exact moment for an action with remarkable precision.
Crucially, this high performance does not come at a high computational cost.
The model delivers this moment-finding capability at a fraction of the compute cost and with 4x the execution speed compared to larger model categories.
This efficiency results in the sub-second latency required for safe and responsive physical interactions in real-world robotics.
Performance Metric Result Significance
Progress Classification Accuracy 57.4% Outperforms previous and competing frontier models in understanding task stages.
Moment-Finding Accuracy 91.3% Achieves high precision in identifying critical moments for action.
Moment-Finding Mean Absolute Distance 0.96s Pinpoints task moments with sub-second temporal accuracy.
Moment-Finding Efficiency 4x faster execution speed Delivers high performance at a fraction of the compute cost of larger models.

Innovations in Tool Orchestration and Reasoning

Beyond raw numbers, Gemini Robotics ER 2 introduces a substantially improved tool orchestration workflow.
This enhancement allows for more seamless and intelligent use of tools in complex sequences.
Internal benchmarks show that it consistently outperforms its predecessor, ER 1.6, across three distinct control modes: real VLA (Vision-Language-Action), simulated VLA, and human tele-operation.
This consistent outperformance highlights the robustness and versatility of the new architecture.
Underpinning these improvements are advancements in the model's core spatial reasoning capabilities, allowing it to better comprehend and interact with the physical layout of its environment and the objects within it.

Advanced Sensing and Multi-modal Understanding

A key breakthrough in ER 2 is its ability to process dynamic visual information for more reliable task evaluation.
Success and failure detection now operates on raw video feeds instead of static snapshots.
This allows the system to catch mid-execution failures in real-time, a critical feature for preventing errors and ensuring task completion.
The model’s sensory perception has also been significantly expanded in its general instrument reading capabilities.
It can now accurately read data from a wide variety of sources, including digital displays, linear scales, rulers, and even liquid thermometers, having been tested across 10 different instrument types.
This is complemented by enhanced spatial VQA (Visual Question Answering), which leverages Gemini's native advancements in multi-modal understanding to answer complex questions about the visual environment.
Across its core functions—including success detection from both images and video, ERQA (Embodied Reasoning Question Answering), and generalized instrument reading—Gemini Robotics ER 2 consistently achieves the highest accuracy, solidifying its position as a new leader in robotic intelligence.


4. Prioritizing Safety: Gemini Robotics ER 2's Robust Safeguards

As Google DeepMind pushes the boundaries of robotic capabilities with the new Gemini Robotics 2, this section connects to the main announcement by focusing on a critical, non-negotiable aspect of its development: safety. The advancements in general intelligence and control are paralleled by equally significant progress in ensuring these robots can operate reliably and securely around humans.

Building Safer Robots for Real-World Deployment

At its core, Gemini Robotics ER 2 is engineered with the explicit goal of making robots safer for deployment in real-world, unpredictable environments.
This design philosophy has resulted in what Google considers its safest robotics model to date.
The model's safety protocols are not just theoretical; they manifest in tangible, observable behaviors designed to protect nearby humans.
A key demonstration of this is its ability to successfully halt a humanoid robot's operation the moment a person comes into close proximity.
Crucially, the system doesn't require manual intervention to restart; it is capable of autonomously resuming its work, but only once it has verified that the area is clear and safe to do so.

Benchmarking Human Interaction and Proximity

The safety improvements are not merely anecdotal but are backed by measurable performance gains on established robotics safety evaluations.
Gemini Robotics ER 2 achieves significant gains on both the Safety Instruction Following and Human Proximity benchmarks, setting a new standard for performance in these areas.
When compared directly, it substantially outperforms its predecessor, ER 1.6, as well as other frontier models, showcasing a clear leap forward in its ability to operate cautiously and predictably around people.
Model Safety Instruction Following Benchmark Performance Human Proximity Benchmark Performance
Gemini Robotics ER 2 Achieves significant gains; outperforms prior models. Achieves significant gains; outperforms prior models.
Gemini Robotics ER 1.6 Outperformed by ER 2. Outperformed by ER 2.
Other Frontier Models Outperformed by ER 2. Outperformed by ER 2.

New Standards for Robotic Safety Evaluation

Recognizing that advanced models require more sophisticated testing, Google has also introduced a new benchmark specifically to evaluate a foundation model's ability to act as a safe Vision-Language-Action (VLA) orchestrator.
This new safety benchmark is comprehensive, designed to rigorously test a model's capacity in four critical domains.
It evaluates the model's ability to consistently enforce safety constraints, actively monitor its environment for potential hazards, accurately assess the physical feasibility of a requested action, and proactively seek human clarification when faced with ambiguity or uncertainty.
For those seeking to understand the technical underpinnings of these advancements, a detailed safety technical report has been made available.


5. Access and Integration: Gemini Robotics ER 2 for Developers

Following Google's July 31st announcement of the Gemini Robotics 2 model, this section details the specific pathways and resources developers can use to access and integrate its advanced capabilities into their robotics projects.

Public and Enterprise Access

Google has made Gemini Robotics ER 2 accessible through multiple channels to cater to both individual developers and large-scale enterprise users.
For general public access, the model is available to developers through the Gemini API and Google AI Studio, allowing for broad experimentation and application development.
For enterprise clients seeking more tailored solutions, Gemini Robotics ER 2 is also available in a private preview on the Gemini Enterprise Agent Platform.
Platform Access Level
Gemini API Public
Google AI Studio Public
Gemini Enterprise Agent Platform Private Preview

Empowering Developer Integration

To accelerate development and ease the integration process, Google has provided a suite of resources.
Code examples are available on Github to help developers get started quickly with practical implementations and common use cases.
A key feature for creating sophisticated robotic agents is the model's direct support for streaming multimodal data.
Developers can stream video, audio, or text directly into the model, which is essential for building an agentic setup that can perceive and react to its environment in real-time.

Flexible Control Interface Options

Gemini Robotics ER 2 provides a high degree of flexibility in how developers command and control robotic hardware.
The platform allows developers to declare low-level control interfaces as tools that the model can invoke.
This means specialized systems, such as other VLA models or dedicated navigation APIs, can be wrapped as a tool, enabling the Gemini model to delegate specific, complex physical tasks to the most appropriate subsystem.


6. The Minds Behind the Machine: Key Contributors to Gemini Robotics ER 2

This section of our analysis on Google's July 31st announcement of Gemini Robotics ER 2 focuses on the key individuals whose expertise was instrumental in this technological leap.
While such a project involves a large team, the official announcement and accompanying documentation highlight the contributions of several key engineering leaders.

Leading Engineers and Researchers

The development of this advanced AI for humanoid control was spearheaded by a talented group within Google's robotics division.
Among the notable contributors is Steven Hansen, who holds the position of Senior Staff Software Engineer.
His senior role indicates significant technical leadership and architectural oversight on the project.
Also credited for his work is Peng Xu, a Staff Software Engineer, whose contributions were crucial to the software's implementation and success.


7. The Road Ahead: Future Directions for Gemini Robotics ER 2

Following the announcement of Gemini Robotics ER 2 on July 31st, Google has outlined a clear vision for the model's evolution, focusing on expanding its capabilities and nurturing the broader robotics field.

Expanding Robotic Capabilities

The immediate future for Gemini Robotics ER 2 centers on pushing the models to tackle progressively more sophisticated challenges.
Having established a new benchmark for full-body control and real-time reasoning, Google's plans now involve training the system for even more complex tasks.
This trajectory suggests moving beyond single-step commands or structured environments towards multi-stage, dynamic problem-solving that requires a deeper level of situational awareness and adaptability from the robot.

Fostering the Robotics Ecosystem

Beyond enhancing the model's own intelligence, a core component of the strategy is to serve a larger purpose within the industry.
Google has stated that its overarching goal is to accelerate the development of genuinely helpful robots.
This indicates a commitment to translating advanced research into practical, real-world applications more quickly.
Furthermore, the company aims to actively support the broader robotics community, positioning Gemini Robotics ER 2 not just as a proprietary technology but as a catalyst that can empower other researchers and developers to innovate and build upon this new foundation.