Skip to content
End-to-End Autonomous Driving Benchmark: How I Test the Homework

End-to-End Autonomous Driving Benchmark: How I Test the Homework

End-to-end autonomous driving benchmark: a repeatable scorecard for latency, safety, trajectory quality, and OTA regressions, built for engineers without PR...

July 29, 2026, 7:10 a.m., Ann Arbor, Michigan. The test vehicle stopped at a clear green light, waited three seconds, then rolled forward as if it had discovered a new traffic rule. Severity: 3 out of 5. Nothing collided, but the decision was unexplained and repeatable. That is exactly why I built an end-to-end autonomous driving benchmark instead of trusting a dashboard full of marketing metrics.

Your car talks. I check his homework. I care less about a claimed sensor count than what the complete system does when a cyclist appears beside a parked truck, a delivery van blocks a lane, or road paint disappears under construction dust. The benchmark records the input, the chosen trajectory, timing, intervention, and whether the behavior makes sense to a human supervisor.

What an end-to-end autonomous driving benchmark should measure

“End-to-end” is often used too loosely. A camera-to-control model is not automatically a complete driving evaluation. A useful end-to-end autonomous driving benchmark follows the chain from scene capture through perception, prediction, planning, control, and vehicle response. If the system brakes late, I want to know whether the cause was a missed object, an incorrect forecast, a cautious planner, or actuator delay.

My harness separates four practical measures. First is safety: collisions, near misses, lane violations, unsafe gaps, and unnecessary emergency braking. Second is trajectory quality: lateral smoothness, speed choice, clearance, and whether the path follows the apparent road rules. Third is responsiveness: sensor-to-decision latency and the time between a new hazard and meaningful control. Fourth is recovery: what happens after the expected route or assumption becomes invalid.

A single composite score hides too much. A system that drives smoothly but misses a stopped vehicle should not average out as “good.” I publish category scores, raw event clips, and a severity rating from 1 to 5. Severity 1 is cosmetic awkwardness. Severity 3 is a decision requiring a supervisor or uncomfortable intervention. Severity 5 is a collision or behavior with immediate serious risk.

Illustration for end-to-end autonomous driving benchmark

The test setup and the uncomfortable details

The test setup matters more than the spreadsheet. I use repeatable routes with ordinary complexity: marked intersections, unprotected turns, parked vehicles, uneven pavement, temporary signs, and mixed traffic. Each run begins with a clean software version, vehicle configuration, weather note, map status, and timestamp. An OTA update gets its own baseline comparison rather than a fresh score with no historical context.

For simulation, I replay recorded sensor sequences and inject controlled changes. A pedestrian can appear one second earlier. A lead vehicle can brake harder. A lane marking can be removed. Those changes do not replace public-road testing, but they help isolate causality. If the car fails only when the object enters from the right edge, that is more useful than a vague statement that the model “struggled with pedestrians.”

I also log false positives. A car that slams the brakes for an empty plastic bag is not safe in the same way as a car that ignores a real obstacle. The first event can produce a rear-end risk; the second can produce a direct impact. Both belong in the end-to-end autonomous driving benchmark, with different severity and different engineering owners.

Repeatability is the hard part. I ran one construction-zone scenario three times and got three different answers because traffic timing changed the planner’s available gap. That result was messy, not useless. It showed that the scenario needed controlled traffic timing and more runs before assigning a regression. Error bars belong in the report, even when they make the chart less impressive.

Comparing systems without worshiping the score

An end-to-end autonomous driving benchmark should compare behaviors, not just vehicles or model labels. Two systems can reach the same destination while making very different safety tradeoffs. One may take a conservative three-second gap; another may merge quickly but smoothly. A scorecard must preserve those distinctions.

For each scenario, I record the expected action before reviewing the video. Then I compare the vehicle’s action against that expectation. Here is the basic structure: what happened, what it should have done, and what the gap means. If a car stops for a pedestrian who is clearly outside its path, I mark unnecessary yielding. If it continues while the pedestrian moves toward the lane, I mark threat estimation. The clip receives a severity rating and a reproducibility note.

Brand names do not change the method. Tesla, Waymo, Mercedes-Benz, General Motors systems, and open-source stacks such as openpilot or Autoware can all be evaluated against the same scenario definitions, provided the comparison states the operating domain. A supervised driver-assistance feature is not equivalent to a geofenced autonomous service. Mixing those categories creates a chart that looks definitive and says very little.

The benchmark also needs a human-intervention ledger. I record why the driver took control, how much time remained, and whether the intervention prevented a likely violation or merely corrected an awkward choice. A high intervention count is a warning, but the reason behind each intervention is the real data.

Visual context for end-to-end autonomous driving benchmark

OTA updates and regression hunting

The most useful run is often the one after an update. An end-to-end autonomous driving benchmark gives me a fixed reference set, so I can ask a narrow question: did version 12.4 change behavior on the left-turn scenario? I do not need a press release to answer that. I need the same route, similar conditions, synchronized logs, and a before-and-after comparison.

Regression hunting starts with behavior changes that look small. The vehicle may now hug the centerline, wait longer behind a bus, or hesitate when a lane splits. A smoother ride can conceal a larger gap in hazard response. I tag each change as improved, degraded, unchanged, or inconclusive. “Inconclusive” is a valid result when lighting, traffic, or localization prevents a fair comparison.

The practical workflow is simple. Save the old software image and calibration. Run the fixed suite. Archive sensor and control logs with hashes. Repeat any surprising event. Review the clips blind when possible, so I am not unconsciously searching for the promised improvement. Then publish the raw observation alongside the interpretation. Readers should be able to disagree with my judgment without guessing how I reached it.

What readers should demand from a benchmark

A credible end-to-end autonomous driving benchmark should publish its scenario definitions, operating limits, scoring rules, and failure examples. “The system handled urban driving” is not a result. “The system completed 18 of 20 defined left-turn trials, with two interventions caused by late cyclist detection” is at least testable, assuming the logs exist.

Ask whether the route was known in advance. Ask whether the driver was coached. Ask how many attempts were discarded and why. Ask whether the car was running production software or a special evaluation build. Those details are not footnotes; they determine what the number means.

My own scorecard will remain intentionally boring: version, route, weather, scenario, outcome, severity, latency, intervention reason, and reproducibility. No leaderboard theater. No “industry-leading” label unless the evidence earns it. The goal is not to declare that autonomous driving is solved. The goal is to find the next stupid decision, reproduce it, and make the gap smaller.

That is the standard I want from every end-to-end autonomous driving benchmark. Show the homework, including the wrong answers.

The Timing Log

0 entries · timing stand

No observations filed for this run yet.

Log an observation

Course observers & crew — file what you saw at the trap. Entries are stamped as witnessed.