AI

Google DeepMind Unveils Gemini Robotics 2 for Whole-Body Control

By Dillip Chowdary July 31, 2026 4 min read
Google DeepMind Unveils Gemini Robotics 2 for Whole-Body Control

Google DeepMind announced Gemini Robotics 2, a foundational Vision-Language-Action (VLA) model designed to provide whole-body physical intelligence for humanoid and mobile manipulator robots. The model bridges the gap between high-level reasoning and millisecond-level torque adjustments across multi-joint kinematic chains.

Unlike modular architectures that separate vision processing from motor controllers, Gemini Robotics 2 operates end-to-end, converting raw camera streams and natural language instructions directly into coordinated motor trajectories. Robotics engineers tuning control scripts can utilize the [Code Formatter](/tools/code-formatter/) for clean code execution.

Tech Pulse Daily

Get tomorrow's tech pulse first

Deeply analytical tech news delivered to your inbox every morning. Free, no spam.

Unifying High-Level Spatial Planning and Low-Level Motor Kinetics

In laboratory demonstrations, robots powered by Gemini Robotics 2 successfully navigated cluttered environments, folded delicate textiles, and operated complex industrial power tools with zero-shot domain adaptation.

Zero-Shot Task Generalization in Dynamic Unstructured Environments

DeepMind's benchmark results show a 4x reduction in task failure rates compared to first-generation models, signaling that humanoid hardware is nearing commercial readiness for warehouse and logistics operations.

Key Takeaway

Google DeepMind introduces Gemini Robotics 2, an end-to-end vision-language-action model enabling real-time whole-body coordination for complex humanoid tasks.