These Tech Workers Made ChatGPT Drive a Toyota Corolla

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase. We do…

By Vane September 29, 2026 5 min read
These Tech Workers Made ChatGPT Drive a Toyota Corolla

Three tech workers in San Francisco connected ChatGPT to a rented Toyota Corolla and made it drive a cone course in a public parking lot. The experiment used a general-purpose chatbot running on a laptop rather than a purpose-built driving system trained on millions of hours of data. The team, who call themselves DrivingBench, has posted their code, prompts, and video footage online. They hooked up GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol to the vehicle and gave the models control over steering, the accelerator, and the brakes. The chatbots then navigated a small obstacle course.

The goal

The objective was to determine if untrained, off-the-shelf frontier AIs can drive a real car. Results were mixed. Grok, Sol, and Fable managed only a few metres before failing to finish. GPT-6 Astra eventually learned to navigate the course and complete it, though significant troubleshooting was required.

The team

DrivingBench consists of Aditya Ramabadran, Tobias Gessler, and Simon Mahns. They met at their day job at Axiom Math, an AI math startup. Ramabadran told 404 Media the idea occurred while the three were at an ice cream shop. They had seen demos of Astra and similar models performing robotic tasks, such as painting or moving objects, which suggested a level of spatial reasoning previously unseen. They asked if there was a way to get LLMs to drive a car.

Why driving a car?

The team wanted to prove it was possible and create a new benchmark for LLMs performing real-world tasks. Ramabadran noted it was not meant to show practicality for daily use. He said it would be shocking for people to see models good enough to drive a vehicle in real life, even at low speeds in a parking lot.

The setup

They rented a Toyota Corolla and sent Gessler to do it. The team did not tell the rental company they planned to hook chatbots to the vehicle. They used Comma, an off-the-shelf system that allows users to install a self-driving kit on unsupported cars. Comma has two cameras pointed at the road that feed data back to the laptop running the LLMs. Gessler said the Toyota Corolla is one of the most popular cars for this kit and the system is open source, making it easy to modify for their needs.

Comma resembles a dashcam. Users attach it to the windshield and it watches the road, recording video and integrating with the car’s electrical systems and sensors. It provides a rudimentary form of self-driving. The National Highway Traffic and Safety Administration announced an investigation into Comma last week after two crashes involving the system killed three people.

How it worked

For DrivingBench, Comma provided an easy way to get the car talking to a chatbot. One person sat in the driver’s seat with a laptop or passenger seat prompting the LLM. Another person pressed a button on the steering wheel to give the command to start. From there, the system monitored the car to ensure the model did not crash into a wall. Mahns said they kept a foot over the brake just in case.

The prompt

DrivingBench’s prompt is on its website and runs fewer than 600 words. It instructed the model to drive through a backwards-U-shaped parking lot and stay between cones. The finish line was a wide parking spot marked by numerous blue mini-cones. Evaluation was based primarily on distance covered without leaving boundaries or collisions, with a secondary objective to complete the course in less time. The rest of the prompt covered specific instructions about the physical limitations of the Corolla and the parking lot.

The problems

Issues began before the car moved an inch. When the LLMs realised they had been prompted to drive a car, most refused. Ramabadran said specifically with Astra, it would refuse in many situations. They had to change the prompt and rename things through hours of iteration to get it to consistently drive. They tried calling it a simulation, which worked some of the time, but the models would see the images and realise they were in a real parking lot. They ended up calling everything a sandbox, which allowed the car to drive consistently without refusal.

Other problems were more pedestrian and caused delays that forced Gessler to extend the rental. Finding a parking lot held them up several times. They visited high schools, churches, and community centers. Mahns said they got kicked out a couple of times. The first place was a church. They commandeered the entire parking lot and an event was starting. The staff asked if they had permission. Mahns said they replied they could pack up right now.

They were also kicked out of an office building parking lot. They took a corner and a security guard asked if they had permission for taking a quarter of the lot. Mahns said they replied not really, so they packed up. Ramabadran added the people who kicked them out were super chill about it.

Software issues

Coding the software also slowed things down. The first time they attempted this, they vibe coded the bridging software between the LLMs and Comma. Mahns said a true and funny story was that they tried to get Astra to one shot some code and it was pure slop. They had to restart from a new design. This pointed to the limits of these frontier models. Mahns said it could not just autonomously do this because it made thousands of lines of slop. Astra’s first attempt at writing the software created a 200,000 line repository.

What it means

Despite scarce parking lots, vibe-coded software, and chatbots fighting the experiment, only one model completed the course. Mahns remains excited about the results. He called it an interesting demonstration of potential emergent capabilities. A model can fail a turn, and then on the next try, it will be able to do that turn and others it has not seen before. He said he is not saying LLMs will put Waymo out of business, just that it is an interesting demonstration of where things are and things are moving fast.

Mahns added these experiments are good because LLMs will increasingly affect things in the real world, not just on screens. Most people familiar with ChatGPT have used it in a white collar, computer-based context. He said it is imminent that this capability will start having more impact in the physical world. Seeing the sparks of this happening and the shift of utility into the physical world is somewhat exciting.

Ramabadran focused on latency and testing high consequence tasks. He noted these models can take like 20 seconds to think and give a response. If you imagine driving at 2mph, which is like one metre per second, and you take 10 seconds to think, you have moved like 10 metres. If you are driving 10 metres blind, that is pretty bad. Ramabadran said it is cool to have a benchmark where latency is part of the benchmark.

Scroll to Top