These Researchers Made AI Drive a Toyota Corolla to Get In-N-Out

Three engineers at the startup Axiom drove a 2024 Toyota Corolla through a drive-thru lane to pick up lunch at In-N-Out Burger.…

By Vane October 7, 2026 3 min read
These Researchers Made AI Drive a Toyota Corolla to Get In-N-Out

Three engineers at the startup Axiom drove a 2024 Toyota Corolla through a drive-thru lane to pick up lunch at In-N-Out Burger. Aditya Ramabadran, Simon Mahns, and Tobias Gessler sat in the passenger and back seats while they asked OpenAI’s GPT-6 Astra to control the vehicle. They connected a laptop to the car’s power steering and linked a chat interface to cameras mounted on the windscreen. A safety driver kept a foot near the brake pedal throughout the trip. The AI model, typically used for text and code, guided the car to the take-out window. One engineer noted that artificial general intelligence might be closer than expected.

From Screen to Street

Self-driving cars usually run on algorithms trained specifically for the task. This experiment used a model designed to output text. The operation happened on the fly without prior coaching. The stunt suggests language-based AI models are gaining a basic understanding of the physical world.

Nothing alarming occurred during the lunch run. The team admits that placing a general-purpose model in charge of a moving vehicle is a high-stakes undertaking. As physical understanding improves, AI models may interact with the real world in new ways that could be dangerous.

Current smart models excel at answering complex questions and performing virtual tasks. Their skills remain limited to computers and the internet. Move these models into the real world and they often become confused. AI companies still view physical reasoning as a frontier they have not conquered.

Andrew Dai, CEO of Elorian AI, previously worked at Google DeepMind. He says better visual reasoning will open new applications. These include systems that understand if diners are enjoying a meal or robots capable of functioning in a home. Dai notes that robotics is a crucial test case for physical reasoning skills. “It’s pretty essential for robotics,” he says. “You can’t really imagine home robotics without this.”

Elorian and Scale AI recently developed a new benchmark called Humanity’s Sixth Sense. This tool measures a model’s ability to understand physical scenes. Xingang Guo, a research scientist at Scale AI involved with the benchmark, explained their goal. “Most visual AI research is about perception,” he says. “We wanted to ask what it would actually take for a model to understand a scene intuitively, the way a person does without thinking about it.”

The engineers conducted the In-N-Out project outside their regular work at Axiom. They wanted to measure how well models operate in the messy real world. The idea came while hanging out one weekend. They noticed models like Astra could build complex 3D simulations and wondered if this translated to navigating the real world.

They live in the Bay Area and are exposed to Teslas and Waymos, which Ramabadran says may have influenced their thinking. The team tried testing SpaceXAI’s Grok before moving to the latest models from OpenAI and Anthropic. At first, the models refused to take control. They responded with messages like, “I can help interpret road images, but I can’t issue motion commands to a physical car.” Careful prompting eventually coaxed them into going further. The engineers were stunned when the models began driving.

The engineers say big AI companies seem focused on improving spatial reasoning in their latest models. They doubt this means actual training to drive cars. The models’ vehicular talents likely materialized as part of training focused on 3D reasoning. Mahns describes this as an emergent capability of just scaling up the multimodality of the model. This refers to inputs like images, video, and 3D models used to train the systems.

A new benchmark from the trio, called DrivingBench, suggests these models have a long way to go before passing a driving test. The benchmark measures a model’s ability to drive around a simple course set up in a parking lot. Only Astra completed the course, and very slowly. Claude Fable 5.1 reached 45 percent of the way around, while Grok made it just 11 percent.

Ramabadran says the latest models seem capable of adjusting to errors. They adapted to the car’s controls and improved their driving in real time. The models appeared to be adjusting or in-context learning based on their mistakes and learning how to better navigate the controls.

What it means

For people making things, this experiment shows that general AI models can handle physical tasks without specific training for those tasks. However, performance remains poor in controlled tests. A model that drives a car to buy lunch is not ready to replace a human driver on public roads.

Scroll to Top