At 9:14 p.m. on a wet Tuesday in Ann Arbor, my test vehicle saw a traffic cone lying across the right lane, moved left to avoid it, and then drifted back toward the cone before correcting again. Nothing crashed. That is not the same as passing. Self-driving car edge case testing exists to expose the gap between a system that completes a drive and one that makes a defensible decision under unusual conditions.
I treat each odd maneuver like a software bug. I save the camera frames, localization data, vehicle speed, planned trajectory, control commands, and system version. Then I ask three questions: what happened, what should have happened, and how large was the gap? Without that record, an edge case becomes a story. With it, the case becomes a regression test.
What counts as an edge case?
An edge case is not necessarily a rare object. It is a situation where the assumptions connecting perception, prediction, planning, and controls become unreliable. A pedestrian partially hidden behind a parked van qualifies. So does a normal sedan stopped beyond a crest, where the planner has only a fraction of a second to react after the vehicle becomes visible.
I also include ordinary scenes with unusual combinations: faded lane markings during light rain, a delivery truck blocking a temporary sign, a cyclist moving beside a right-turning vehicle, or a police officer gesturing while the traffic signal remains red. The individual ingredients are familiar. Their timing and interaction create the test.
My severity scale runs from 1 to 5. Severity 1 is an awkward but safe response, such as unnecessary braking. Severity 3 is a maneuver that requires a human takeover or creates uncomfortable proximity. Severity 5 means a collision, a near collision, or a decision that could reasonably produce one. The rating is about behavior, not how dramatic the video looks.

Building a useful test harness
Good self-driving car edge case testing starts with repeatability. My harness stores each scenario as a bundle: map location or simulator seed, weather and lighting notes, actor positions, initial vehicle state, software build, and expected behavior. A replay should produce the same inputs even when the vehicle stack is updated months later.
For a real vehicle, I use a clearly defined route, a trained safety driver, and a controlled site whenever the scenario involves an actor or obstruction. I never create a live-road test that depends on another road user reacting correctly. Simulation is better for dangerous combinations, but simulation is not automatically truth. A simulated pedestrian with perfect physics can hide the messiness of a real person, a loose jacket, or a partially occluded body.
The harness records more than a pass or fail. I capture time to first detection, classification confidence when available, predicted path changes, maximum lateral error, braking onset, minimum distance, and whether the final trajectory was stable. Latency matters. A vehicle that recognizes a child but waits too long to brake has not solved the scenario.
The cases that break the stack
The most valuable self-driving car edge case testing targets interfaces between modules. Perception may correctly identify a road worker, while prediction assigns an implausibly low probability to the worker stepping into the lane. Planning then produces a trajectory that looks smooth but leaves no practical escape margin. Controls follows the trajectory exactly. Each component can appear reasonable in isolation while the complete vehicle behaves badly.
Occlusion is a recurring offender. Put a pedestrian near a parked SUV and vary only the visible body area. At full visibility, the system should slow and pass with space. At partial visibility, it should preserve enough margin for the hidden person to emerge. I record the exact point where behavior changes, not just whether the final result was safe.
Another useful family involves ambiguous road geometry. A freshly paved section can have no visible lane lines, while old markings remain at the shoulder. A planner that anchors too strongly to paint may place the car incorrectly. I test the same route with clear lines, faded lines, conflicting lines, and temporary construction markings. The expected response is not always a single path; it is often a controlled reduction in speed and a conservative position.
A concrete replay example
Here is a representative test from my log. A stopped vehicle occupied a travel lane just beyond a hill, and an approaching car was visible in the opposite lane. The system detected the stopped vehicle late, began a lane change, then canceled the maneuver after detecting the oncoming car. It returned toward the stopped vehicle before braking firmly.
The expected behavior was earlier deceleration, followed by a wait behind the obstruction until the opposing lane was clearly available. The gap was not a collision, but the response had two unstable commitments: first to pass, then to abort without enough space. I rated it Severity 3 because a human driver would need to intervene in a narrow timing window.
This is where self-driving car edge case testing earns its keep. A single video review might label the event “hard braking.” The replay shows a planning problem, a timing problem, and a missing fallback state. I can rerun it after an update and determine whether the system now slows earlier, waits longer, or simply fails in a different way.

Measuring OTA updates without fooling yourself
An over-the-air update should be evaluated against a fixed baseline, not against memory. I run unchanged scenarios before and after the update, then add a small set of new cases designed around the release notes. If a manufacturer says it improved construction-zone handling, I test barrels, temporary arrows, lane shifts, workers, and contradictory paint rather than accepting the label.
I also watch for regressions. A change that improves pedestrian detection can increase unnecessary braking near signs or shadows. A more cautious planner can reduce collision risk while creating hesitation at unprotected turns. Those are engineering tradeoffs, not reasons to declare victory from one successful drive.
The output is a scorecard with scenario-level results, severity ratings, and raw clips. I keep inconclusive outcomes visible. I have run the same case three times and received three different trajectories because small localization differences changed the planner's available options. That is not a clean pass or fail; it is evidence that the scenario needs more controlled variation.
What a credible result looks like
Credible self-driving car edge case testing names the system version, hardware configuration, route conditions, safety protocol, and scoring rule. It separates observed facts from interpretation. “Vehicle speed fell from 25 to 12 miles per hour” is an observation. “The planner became cautious” is an interpretation that needs supporting evidence.
I do not treat a smooth video as proof of competence, and I do not treat an ugly maneuver as proof that an entire system is unsafe. The useful unit is the reproducible case. If the result survives replay, controlled variation, and an OTA comparison, it belongs in the benchmark. If it cannot be reproduced, I keep it in the log with that limitation stated plainly.
That is the standard I use for every release: a repeatable input, an explicit expected response, measurable trajectory behavior, and a severity rating when the system misses. Your car talks. I check his homework.