Elon Musk told shareholders in July that Tesla’s Optimus robot will eventually possess human and then superhuman dexterity. The company CEO predicts the machine could automate nearly all human labour, from hauling sheet metal to folding laundry, for as little as $20,000 each. Speaking at the World Economic Forum in Davos, Switzerland, in January, he added that public sales might begin by the end of 2027.
In this article
A white robot with a black head and torso has appeared on video feeds performing tasks like dancing, putting trash in bins, and pressing microwave buttons. Sometimes it falls backward while handing out water bottles or struggles to iron a shirt. This is Tesla’s Optimus, a humanoid robot that Musk believes could become the biggest product in history.
Marc Andreessen, cofounder of Andreessen Horowitz, has stated that robotics could become the biggest industry in the history of the planet. Jensen Huang, CEO of Nvidia, said humanoid robots would match human-level ability this year. According to Morgan Stanley, the number of robots resembling and acting like humans is likely to reach nearly 1 billion by 2050, creating a market worth over $5 trillion.
These claims rely on the idea that the same AI advances behind tools like OpenAI’s ChatGPT and Anthropic’s Claude will enable robots to imitate human movement. Many robotics researchers are skeptical, arguing that assumptions about using intelligence built on language and images to master the physical world underestimate the challenges. Yann LeCun, often referred to as one of the godfathers of AI, said at the January Davos conference that none of the companies building humanoid robots has any idea how to make those robots smart enough to be useful.
Researchers also point out that conflating humanoid robots made to resemble people with generalist machines able to learn and perform multiple tasks is misleading. Jonathan Hurst, cofounder and chief robot officer of Agility Robotics and professor of robotics at Oregon State University, explains that it is very easy to make a robot that looks like a person. It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.
Tensions over whether all-purpose humanoid robots are just around the corner or nowhere in sight are playing out in robotics labs across the country. Hype over timelines is obscuring painstaking but meaningful progress.
A decade or so ago, a series of breakthroughs led to a generative AI revolution that turned the long-imagined possibility of artificial intelligence into reality. Roboticists believe an equally transformative revolution is possible in robotics, one that will endow machines with physical intuition and fluidity that has long been out of reach. As progress inches forward, the question is whether the same methods and tools that fuelled advances in AI are enough to get there, or if an entirely new path is required.
Robots meet advanced AI
To see one of the smartest robot brains working today, it is worth looking at what Google DeepMind can do with a piece of equipment called ALOHA 2. The name stands for “A Low-cost Open-source Hardware System for Bimanual Teleoperation.”
Roboticists have long clashed over whether a humanlike form is necessary for generalist robots. Proponents argue it will help them slot into the world as it exists. Detractors say it is not worth the trouble. ALOHA 2 reflects the second way of thinking. It is not much to look at. It is just a pair of arms, some grippers, and a couple of cameras. Despite its seeming simplicity, it is a workhorse for researchers at Google DeepMind who use it to test their most advanced AI for robotics system, Gemini Robotics, in their various labs.
When controlled by Gemini Robotics, ALOHA 2 becomes a generalist robot. It can perform any number of tasks based on examples it has been trained on. Ask it to pack a lunchbox and, as evidenced by a video of this exercise, it can use two pincer grippers to delicately place a piece of white bread into a Ziploc bag. It closes the bag, places a bunch of grapes in a Tupperware container, secures the lid, and then carefully moves the items into a lunchbox before zipping it up.
The result is not a great lunch. The fact that the robot can put it together represents an objective step forward from what was possible even three years ago.
This is in large part due to AI and its impact on what are known as robot policies. These controls how a general-purpose robot assesses and understands its surroundings, plans how to move within them, and then performs its task correctly.
Historically, these policies were based on rules developed by engineers who hard-coded them into the robot’s software. Thousands of lines of code would determine each millimetre of a robot’s movements in hundreds of tasks. What has been happening for the last few years—and what is largely responsible for the optimism about generalist robots—is that robot policies are being handed over to advanced AI systems instead of being coded into the robot’s software.
This first happened with VLMs, or vision-language models. These are similar to large language models but are trained on images as well as words. Show a VLM a picture of a coffee spill and ask it to find a tool to clean up the mess, and it can identify a nearby cloth. This sort of immediate contextual understanding did not exist a couple of years ago when robot policies were hard-coded.
Next came vision-language-action models. These enable robots to assess their environment and take action within it. The models do this by adding yet another component: motion commands. VLAs are trained on a series of images or videos related to performing a given task along with associated data about how a robot arm moves to perform it. That movement data is typically collected through teleoperation, in which a human uses remote controls to lead a robot through an action. This sort of training allows the AI to learn how to command the robot to move and operate during a given task.
Place a VLA-powered robot in front of a desk and tell it to “close a laptop” or “wrap up the headphone wire,” and it will survey the scene, identify the relevant object, plan a way to execute the request, and then swing its arms into action—at least if it has seen this task accomplished before.
The Gemini Robotics model is a VLA, trained on many hours of human demonstrations depicting a vast array of different actions. As a result, it can perform relatively complex tasks like picking up snow peas with kitchen tongs, doing origami, or putting together a simple lunch. It is impressive, but there is a glaring limitation. For now, if a robot controlled by a VLA is asked to perform a task that falls outside its training set, it is highly likely to fail.
“Thinking about the space of all tasks, a real generalist policy would be able to do everything along that spectrum,” says Edward Johns, a robotics professor at Imperial College London. Today, though, a Gemini Robotics model can do only “a few things here and a few things there.”
The search for true generality
So how do we get robots to be able to do more things? The usual answer is probably not surprising: Train them on more data.
More data, the thinking goes, equals more examples, and more examples equals more generality. Google DeepMind, for instance, wants to pull together “as much data as possible,” says Pannag Sanketi, a former tech lead in robotics at the company who is currently working on his own AI robotics project. But where to get it? Large language models had the benefit of oceans of existing text for training. There is no corresponding pool of high-quality physical demonstrations on which to train robots.
Researchers have a few ways to make up for this, but all have flaws. One is to employ large numbers of people to create and collect teleoperation data. This is costly and time-consuming. Another is to train VLAs on videos of people performing activities. The resulting data quality is poor. Yet another is to deploy robots in the real world and use data collected from those deployments.
What it means
For people making things, the immediate reality is that robots are becoming better at specific, repetitive tasks like packing lunchboxes or placing grapes in containers. However, the gap between what a robot can do when shown a task it has seen before and what is required when faced with a new situation remains wide. Users should expect machines to fail if asked to perform actions outside their training sets, meaning current systems are not yet ready for the unpredictable variety of work found in most homes and factories.



![Rare event prediction on time series that change structure mid-stream? [D]](https://ai-maestro.online/wp-content/uploads/2026/05/rare-event-prediction-on-time-series-that-change-structure-m-768x432.jpg)