At 6:40 p.m. on a wet Tuesday in Ann Arbor, I gave the same lane-change scenario to three driving systems and got three different answers. One slowed early, one crossed the lane marker, and one treated a shadow beside the curb as a moving object. These are autonomous driving hallucination examples in the useful engineering sense: cases where a system builds an implausible interpretation of the scene and acts on it with too much confidence.
I use the word carefully. This is not a claim that a car has a mind, imagination, or an LLM hidden behind the steering wheel. It is a shorthand for a perception, prediction, or planning output that is unsupported by available sensor evidence. My test harness records camera frames, vehicle speed, map context, planned trajectory, control actions, and driver intervention. The goal is not to collect viral clips. It is to answer three questions: What happened? What should have happened? Where is the gap?
What Counts as a Driving Hallucination?
A bad drive is not automatically a hallucination. A vehicle that brakes for a real, partially obscured pedestrian may be conservative, but it is not inventing an object. A vehicle that stops because a plastic bag crosses the road has made a classification mistake. I reserve the hallucination label for a stronger failure: the system represents something that is not supported by the scene, preserves that representation long enough to affect planning, or produces an action inconsistent with its own available context.
I rate each event from 1 to 5. Severity 1 is a harmless hesitation with no meaningful safety consequence. Severity 2 is an unnecessary slowdown or awkward path choice. Severity 3 requires a driver correction or creates a clear conflict with normal traffic. Severity 4 creates an immediate collision risk or violates a predictable right of way. Severity 5 is a crash or a failure that would be unacceptable even with an attentive safety driver.
That scale matters because a phantom cone and a phantom pedestrian are not equivalent. Both are false objects. Only one normally changes the risk envelope dramatically.

Example One: The Phantom Object at the Curb
Timestamp: 2026-05-18, 7:12 p.m., suburban arterial, light rain. The test vehicle approached a curbside parking lane containing a dark trash bag, a signpost, and a broken reflection from a storefront window. The planner reduced speed from roughly 25 mph to 9 mph, then held the lower speed for several seconds after the object was behind the vehicle.
Here is what happened. The perception output alternated between “debris,” “pedestrian,” and “unknown obstacle” across consecutive frames. Here is what it should have done. It should have maintained a stable uncertainty estimate, selected a safe lateral path, and resumed speed once the object passed. Here is the gap. The system was not merely cautious; it lacked temporal consistency.
I scored this event a 2 because no vehicle or person was endangered. Still, it matters. Repeated false positives teach a driver that braking events are often meaningless. That is a trust failure, and trust failure is a controls problem even when the immediate maneuver is safe. I reran the scene three times. The system hallucinated a person twice and ignored the bag once, which is precisely why one clip is not a benchmark.
Example Two: The Invisible Lane Boundary
A second class appears when the road geometry is visually ambiguous. Fresh asphalt, faded paint, construction seams, and temporary markings can cause a model to infer a lane boundary that is not there. In one controlled simulation, I placed a vehicle on a two-lane road with an old white stripe visible beneath a newer surface. The planner generated a trajectory that hugged the phantom stripe and moved the vehicle toward the centerline.
This is one of the more useful autonomous driving hallucination examples because the error crosses subsystem boundaries. Perception supplied a plausible but incorrect marking. Localization did not reject it. Prediction assumed nearby traffic would remain in its lane. Planning then produced a smooth trajectory that looked reasonable in isolation. Controls executed it cleanly. Every module behaved normally relative to a bad premise.
The correct response is not simply “improve the camera model.” I would add temporal lane evidence, map disagreement checks, and a planner penalty for unexplained centerline proximity. A system should also degrade gracefully: when markings disagree with road edges and vehicle history, it should slow down and center itself rather than confidently choosing one interpretation.
Example Three: The Phantom Yield or Stop Sign
Traffic-control hallucinations are easier to miss because the resulting behavior can look polite. A vehicle sees a rectangular advertisement, a rear-facing sign, or a tree branch at the edge of the image and assigns it stop or yield semantics. It then brakes at an uncontrolled location, waits for a nonexistent crossing conflict, or refuses a permitted turn.
In my log, these events receive a severity 2 when they cause delay and a severity 3 when they create a rear-end risk. The key diagnostic is the mismatch between sign detection and sign orientation. A forward-facing regulatory sign should agree with lane geometry, mounting position, and the direction of travel. A classifier that says “stop sign” while those context checks disagree is not finished making its decision.
This is also where production-system comparisons become dangerous. A software update can reduce false stops in one lighting condition while increasing them at dusk. I do not publish a score from one route and call it a universal ranking. I save the raw frames, repeat the route, and report whether the behavior survived changes in time, weather, and traffic.

Example Four: Predicting a Vehicle That Is Not There
Prediction models can hallucinate motion even when perception is mostly correct. Consider a parked car partly hidden behind a hedge. The system identifies the visible front corner, estimates a likely trajectory, and then behaves as if the parked vehicle is pulling into traffic. The resulting brake may be safe, but the model has promoted a possibility into a near-term event.
The reverse failure is more serious: a moving vehicle is briefly occluded, and the predictor assumes it stopped. The ego vehicle proceeds into a gap that no longer exists. For testing, I record not only the selected path but the predicted occupancy over the next few seconds. A prediction that is wrong once is expected. A prediction that collapses sharply after every occlusion is a repeatable weakness.
The fix is not to eliminate prediction uncertainty. That would be impossible. The fix is to expose uncertainty to planning. A planner should treat an occluded vehicle as a bounded risk, not as either definitely moving or definitely stationary. Calibration curves, replay tests, and counterfactual trajectories tell us more than a single average accuracy number.
How I Reproduce Autonomous Driving Hallucination Examples
My test harness starts with a fixed scenario definition: road layout, weather, lighting, actors, sensor configuration, and expected safe behavior. Each run gets a scenario ID and a software version. I capture the model outputs at a consistent interval, synchronize them with vehicle telemetry, and mark the first frame where the unsupported interpretation appears.
Then I change one variable at a time. I move the trash bag by a few feet, alter the sun angle, add a cyclist, or change the lead vehicle's speed. If the result disappears, that is useful information, not a failed test. It tells me the trigger is narrow. If it persists across five variations, it belongs in the regression suite.
I also separate safety-driver intervention from system intent. A driver touching the wheel does not prove the planner would have crashed, just as a clean-looking trajectory does not prove the perception state was correct. For every entry, I publish the evidence, the severity rating, the repeat count, and the unresolved ambiguity. Your car talks. I check his homework.
What Owners and Engineers Should Watch After an OTA
After an over-the-air update, do not rely on a smoother demo drive. Watch for changes in hesitation, phantom braking, lane centering near construction, and behavior around occluded vehicles. Run the same route before and after the update, using comparable weather and traffic when possible. Save dashboard video, timestamped intervention notes, and the software version displayed by the vehicle.
Engineers should go one step further: retain difficult cases instead of deleting them after a model improves. A regression corpus needs the ugly scenes—the plastic bag, faded stripe, rear-facing sign, and half-hidden car. Those cases expose whether the system learned a robust representation or memorized the appearance of one test.
The practical conclusion from these autonomous driving hallucination examples is not that autonomous driving is useless. It is that confidence must be earned at the scenario level. A car can perform thousands of ordinary miles and still fail on one unsupported assumption. The benchmark is the habit of finding that assumption before the road does.