Intelligent whole-body control with Gemini Robotics 2
Google DeepMind · 2026-07-30 · official · 99,763 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This video is a demonstration by Google DeepMind showcasing "Gemini Robotics 2" running on an Apptronik Apollo humanoid robot. It is presented by Jie Tan, Principal Research Scientist and Director at Google DeepMind, who explains the integration of embodied reasoning and vision-language-action (VLA) models for intelligent whole-body control.
What is shown
- [00:00] Apollo humanoid robot performing whole-body calibration and autonomous walking movements (labeled "Autonomous 1x").
- [00:27] Jie Tan instructs Apollo through a microphone to pack bags for children going to play sports.
- [00:32] An on-screen UI shows a calendar entry: Jessie has a pickleball match at 2:00 PM and Jeremy has a baseball game at 4:00 PM; Apollo parses the schedule and confirms verbally.
- [00:46] First-person and third-person camera views showing Apollo locating baseball gloves, baseballs, and pickleball gear on cluttered storage shelves.
- [00:51] Apollo grasps a baseball glove and balls and places them inside a designated sports tote bag.
- [01:06] Split-screen demonstration of Apollo balancing dynamically on the spot while adjusting its legs and center of mass.
- [01:33] Failure recovery: Apollo misses picking up a pickleball, visually recognizes the dropped ball/failure, and successfully re-attempts grasping it.
- [01:43] Apollo retrieves a pickleball paddle and packs it into the bag.
- [02:04] Jie Tan assigns a follow-up challenge: locating and lifting a tote bag placed on the floor to the robot's left onto a table.
- [02:11] Stress testing: a researcher uses a pole to nudge and perturb the bag on the floor while Apollo dynamically adjusts its stance, squats down, maintains balance, picks up the bag, and stands up.
- [02:38] DeepMind website link displayed (
deepmind.google/gemini-robotics) along with the Gemini Robotics 2 title card.
Claims & numbers
- The video states all robot footage shown is "fully autonomous with Gemini Robotics 2" running at "Real-time footage" ("Autonomous 1x").
- Jie Tan claims the Gemini Robotics embodied reasoning model interprets the environment, vision, and natural language instructions, and then calls a VLA (Vision-Language-Action) model to generate actions.
- Jie Tan states that maintaining balance requires coordinating all actuators from feet to fingertips, and balance adjustments must execute within a fraction of a second.
Notable quotes
- [00:40] Jie Tan: "The Gemini Robotics embodied reasoning model can understand the world, understand what it sees, understand the natural language instructions..."
- [01:03] Jie Tan: "...the robot need to coordinate all the joints and the actuators from feet to fingertip while staying, maintaining balance."
- [02:22] Jie Tan: "Whole-body control is a necessity to achieve that goal."
Assessment
This is an official demonstration video produced by Google DeepMind showcasing Gemini Robotics 2 in a lab environment. The footage is presented as real-time and fully autonomous, highlighting multi-step reasoning, dynamic whole-body balance, and automated error recovery under physical perturbation.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.