TL;DR
Get garage and car supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Researchers tested general-purpose AI models on a short, cone-marked course using a Toyota Corolla, Comma vision hardware and OpenPilot. In 11 runs across four agents, eight failed to complete more than 11% of the course; GPT-6 Astra was the only agent to finish, on its second attempt. The small test does not establish that these models can drive safely on roads.
The researchers—Aditya Ramabadran, Simon Mahns and Tobias Gessler—designed the course as a short point-to-point route in a parking lot. Small cones marked its boundaries and bends. Each agent received the same instructions and was asked to report what it saw and did after each step. The agents were given three attempts, according to The Drive’s account of the test.
The setup connected the models to the car through Comma vision hardware and OpenPilot. Commands were sent over a wireless hotspot, with the agents pausing regularly while responses came back. The interactions took place in a continuous chat, allowing viewers to see the video and the agents’ accompanying messages during each run.
The reported failures varied. Grok mistook a gap between boundary cones for a gate and drove off the course after two commands in its first attempt. Another GPT model reportedly invented a pattern in the cone colors despite the written instructions warning that the colors did not follow that pattern. Several agents did not turn far enough to follow the bends. GPT-6 Astra completed the route on attempt two, though its path reportedly came close to leaving the course near the finish. DrivingBench put the cost of that successful run at $7.74, nearly four times the cost of Astra’s previous attempt, which covered about half as much of the route.
Why a Parking-Lot Test Matters
The experiment highlights a gap between the abilities of general-purpose AI assistants and the demands of controlling a vehicle. A model may describe a scene or follow a conversation, but driving requires it to interpret visual information, make timely decisions and adjust steering in a physical environment where mistakes can have immediate consequences. The reported confusion over cones and steering illustrates those challenges in a limited setting.
That does not show that the tested systems are suitable—or unsuitable in every form—for automotive use. The trial involved a short course, particular hardware and a small number of runs. Still, its results are a reason not to treat broad language or reasoning ability as evidence of dependable vehicle control. One successful course completion is not a safety record, and the test does not establish how an agent would perform in traffic, poor weather or other complex conditions.
As an affiliate, we earn on qualifying purchases.
How DrivingBench Tested the Models
DrivingBench was presented as a way to test frontier, general-purpose models rather than AI systems designed specifically for driving. The distinction matters: the researchers were asking whether models normally used as conversational agents could control a car through a vision-and-control setup, not evaluating a commercial driver-assistance feature or a purpose-built autonomous-driving stack.
The source account describes 11 runs across four agents, but does not give a full breakdown of every model and attempt in its summary. Results included repeated early failures and one successful run by GPT-6 Astra. The researchers’ course and step-by-step chat format made the behavior visible, but the reported sample remains too small to support broad claims about model performance beyond this test.
“Astra’s successful run cost $7.74.”
— DrivingBench, as reported by The Drive
Limits of the Reported Trial
The source does not provide enough detail to determine how the agents would perform across a larger set of courses, repeated trials or different operating conditions. It also does not establish whether the models had been tuned for the task, how their performance compares with a human driver on the same course, or whether the results can be reproduced independently. The specific date of the research report is not included in the supplied material.
Most important, completing a cone-marked parking-lot route does not demonstrate safe public-road driving. The reported results concern a controlled course and a particular hardware setup; there is no evidence in the source that any tested model was approved or being proposed for unsupervised use on public roads.
More Testing Needed Before Road Claims
The supplied report does not announce a next testing date or a planned follow-up. Further evaluation would need to test more runs and routes, report model and setup details, and compare performance consistently before readers could draw wider conclusions about the agents’ driving ability. Until then, the clearest conclusion is limited: GPT-6 Astra finished one test course, while most reported runs failed early, and the trial offers no basis for treating a general-purpose chatbot as a road-ready driver.
Key Questions
Did ChatGPT successfully drive the course?
GPT-6 Astra, identified in the report as the only agent to finish, completed the course on its second attempt. That was one result in a small test, not proof of reliable driving ability.
What happened when Grok tried the course?
According to The Drive’s account, Grok interpreted a gap between boundary cones as a gate and drove off the course in its first attempt, which ended after two commands.
How many tests failed early?
The report says eight of 11 runs across four agents failed to complete more than 11% of the course. The summary does not provide a complete model-by-model results table.
Was this a test of self-driving cars on public roads?
No. DrivingBench used a short, cone-marked parking-lot course and a Toyota Corolla connected to Comma vision hardware and OpenPilot. The reported test does not establish performance or safety in public-road conditions.
Does the result mean GPT-6 Astra is safe to drive a car?
No. Completing one controlled route is not a safety certification or evidence that the model can handle traffic. The source reports no approval for road use and leaves broader performance untested.
Source: rss
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
