Three AI engineers recently commanded OpenAI's GPT-6 Astra—a model built to produce text and code—to pilot a 2024 Toyota Corolla through an In-N-Out Burger drive-thru lane in the Bay Area, according to a newsletter published in Wired. Aditya Ramabadran, Simon Mahns, and Tobias Gessler, engineers at startup Axiom, connected a chat interface to windscreen cameras and the car's power steering, then watched as the language model successfully navigated to the take-out window with no advance instruction. The experiment suggests that AI models originally designed for virtual tasks are starting to develop basic comprehension of physical environments.

The trio's new benchmark, DrivingBench, reveals these models still struggle with real-world vehicle control. Only Astra managed to complete a simple parking lot course, and it moved very slowly. Claude Fable 5.1 finished 45 percent of the route, while Grok covered just 11 percent of the distance. When the engineers first attempted the tests, models refused commands with responses like "I can help interpret road images, but I can't issue motion commands to a physical car," but careful prompting eventually persuaded them to take control.

According to Ramabadran, the newest models appeared to correct their own errors and refine their driving in real time. "The models really did seem to be adjusting or in-context learning based on their mistakes and learning how to better navigate the controls," he tells Wired. The engineers suspect big AI companies are prioritizing spatial reasoning improvements in their latest releases, though they doubt the firms are actually training models to operate vehicles. Mahns describes the driving ability as potentially "an emergent capability of just scaling up the multimodality of the model," referencing the images, video, and 3D models now fed into training systems.

Physical reasoning represents a key challenge AI companies have yet to solve, even as some claim artificial general intelligence has arrived. Andrew Dai, CEO of startup Elorian AI and a former Google DeepMind researcher, says stronger visual reasoning will unlock applications ranging from systems that detect whether restaurant diners are satisfied to robots capable of functioning inside homes. Elorian and Scale AI recently built a new benchmark called Humanity's Sixth Sense to measure how well models grasp physical scenes. Xingang Guo, a research scientist at Scale AI involved in developing the test, explains that most visual AI research focuses on perception, but the team wanted to assess whether a model could "understand a scene intuitively, the way a person does without thinking about it." The engineers' vehicular abilities likely emerged from training aimed at 3D reasoning rather than explicit driving instruction, the trio believes.

Though nothing overtly dangerous occurred during the lunchtime drive-thru run, the engineers acknowledge that handing control of a fast-moving, two-ton vehicle to a general-purpose model carries serious risks. As physical understanding advances, AI models could reach into the real world in both impressive and potentially hazardous ways. For companies evaluating AI deployment in physical environments, the gap between controlled benchmarks and unpredictable real-world scenarios will demand careful governance structures that current product frameworks may not support.