Google DeepMind has released Gemini Robotics 2, an artificial intelligence system capable of directing humanoid robots to perform dextrous tasks such as screwing in lightbulbs and tying trash bags.
In this article
How the system works
The new system combines several models into a single unit. A vision language model, which processes images and video, handles communication with humans and reasons about task execution. Two vision language action models manage physical movement. These models control full-body motion as well as the movement of grippers or hands.
Video demonstrations released ahead of the launch show various robots performing complex actions autonomously. In one instance, Apptronik’s Apollo 2 robot, equipped with hands from Sharpa, organised shelves. Google DeepMind trained the model using human teleoperation, video examples, and simulations. Specific training is still required for AI models to execute a wide range of complex tasks.
While Anthropic and OpenAI have led development in chatbots and coding tools, Google has a stronger history in robotics research. The search giant has previously published significant work on using AI to train robots for useful actions. This release indicates a belief that artificial intelligence must move beyond the digital environment to reach its full potential. The company previously partnered with Boston Dynamics to provide the control systems for legged robots.
What the team says
Carolina Parada, head of robotics at Google DeepMind, describes the release as another milestone toward physical AGI. Her definition is simple: getting a robot to do anything a human can.
Risks and safety measures
Allowing frontier AI models to control robots in workplaces or homes introduces specific risks. Previous research indicates that using advanced AI to control robots can result in unexpected and sometimes dangerous behaviour. Concerns about sudden or unwanted actions in the digital realm emerged recently when an unreleased OpenAI agent hacked several systems.
Parada notes that safety becomes even more critical when robots operate in varied situations. Uncertainty will inevitably appear, requiring a deeper understanding of safety questions.
Google employs a multi-layered safety approach, applying guardrails to each model layer. The company is also introducing ASIMOV-Agentic, a new benchmark for measuring the safety of AI systems collaborating to control a robot. This tool detects whether a command will lead to harmful or uncertain outcomes.
Demis Hassabis, the company’s CEO, has previously stated his hope to develop an AI operating system for multiple robots, similar to the Android operating system found on smartphones.
What it means
For people making things or managing facilities, this shift means robots can now handle physical tasks with a level of autonomy previously reserved for humans. However, it also means operators must rely on new safety benchmarks to ensure these machines do not act unpredictably in the real world.




