At 7:42 on a wet Tuesday morning in Ann Arbor, a test vehicle slows for a delivery van, hesitates, then hands control back to the safety driver. That single event can become a row in a regulatory spreadsheet. It cannot, by itself, tell you whether the planner failed, the perception stack was uncertain, or the driver simply preferred to take over. Good av disengagement report analysis starts by refusing to pretend that one row is a complete safety result.
I have spent enough time reading logs to know the dangerous shortcut: sort companies by disengagement count, declare a winner, and move on. The California Department of Motor Vehicles reports are useful, but they are self-reported operational data with uneven context. The job is to reconstruct the denominator, understand the intervention, and separate a genuine system weakness from a route, policy, or reporting artifact. Your car talks. I check his homework.
What a disengagement actually measures
A disengagement is generally an instance in which the autonomous system is deactivated or control is taken over by the safety driver. The exact reporting rules and vehicle programs matter. A public-road test fleet, a driverless service, and a development vehicle do not produce interchangeable observations.
That distinction is the first filter in av disengagement report analysis. A vehicle traveling 1,000 miles on a constrained freeway has less exposure to unprotected left turns than one running 1,000 miles through dense downtown traffic. Comparing the raw totals without route mix is like comparing two perception models using different camera resolutions.
I record five fields before interpreting any event: date, location type, miles or operating time, whether the intervention was planned, and the reported cause. I also note whether the system was in autonomous mode, whether a safety driver was present, and whether the event involved a collision risk. Missing fields are not minor paperwork problems. They change what the number means.
A low count can reflect strong planning. It can also reflect conservative operating boundaries, limited miles, or frequent pauses that never enter the report as disengagements. A high count can indicate weakness, or simply a fleet operating in harder conditions. The raw number is the beginning of the investigation, not the conclusion.

Normalize the denominator before ranking companies
The most basic calculation is disengagements per 1,000 autonomous miles. If Fleet A reports 20 events over 100,000 miles and Fleet B reports 12 over 20,000 miles, Fleet B has the larger normalized rate even though its headline total is lower. That is elementary arithmetic, but public commentary routinely skips it.
Mileage alone is still incomplete. I separate freeway, arterial, residential, construction-zone, and high-density urban driving when the source provides enough detail. Time of day and weather are useful secondary fields. A system that drives fewer miles in rain has not necessarily solved rain; it may be avoiding it.
For av disengagement report analysis, I use a simple exposure table rather than a single league table. The table has total autonomous miles, total interventions, rate per 1,000 miles, route categories, and the percentage of events with a stated technical cause. If a company reports 80 percent of its miles in easy conditions while another reports 45 percent, the rates need a warning label.
Uncertainty matters too. Small fleets produce noisy rates. Ten events over 5,000 miles can swing dramatically after one unusual intersection, while 100 events over a million miles are more stable. I do not claim that a rate with two reported events is precise to several decimal places. False precision is still false information, even when a spreadsheet produces it automatically.
Read the cause, not just the count
The cause description is where the engineering signal usually lives. Categories such as perception, prediction, planning, localization, hardware, and driver behavior should not be treated as interchangeable. A planner braking for a static object is a different bug from a camera blackout, and both differ from a safety driver taking over because the vehicle is behaving conservatively.
In my av disengagement report analysis, I assign a severity from 1 to 5. Severity 1 is a nuisance intervention with a reasonable fallback. Severity 3 means the behavior was materially wrong but recoverable. Severity 5 means the system created an immediate collision risk, entered an unsafe trajectory, or failed to respond to a critical road user.
This scale is not an official regulatory score. It is a comparison tool. A fleet with fewer events but several severity-5 cases deserves more attention than a fleet with many severity-1 interventions caused by cautious yielding. The event description must support the rating; if it does not, I mark the case ambiguous instead of inventing certainty.
I also look for repeated signatures. Five reports mentioning unusual stops near temporary lane shifts suggest a pattern worth reproducing. Five unrelated hardware resets may belong to one recall or software release. Grouping by symptom, map region, weather, and software version often reveals more than grouping by manufacturer.

Separate regulatory data from a real benchmark
California DMV filings are valuable because they provide a public window into testing activity, but they are not a standardized benchmark of autonomous driving competence. Companies can differ in definitions, reporting detail, fleet composition, and operational design domain. Some records are terse enough that two engineers could reasonably reach different conclusions.
That is why av disengagement report analysis should be paired with controlled testing. Put the same scenario in front of multiple systems when possible: a blocked lane, an ambiguous merge, a pedestrian near the curb, or a temporary traffic-control setup. Record the route, weather, software build, intervention distance, and the exact action the car took.
A repeatable test harness should preserve raw video and vehicle signals, not just a pass or fail label. I want timestamps for detection, braking, steering, and takeover; a map snapshot; and a short explanation of the expected trajectory. If the result changes after an OTA update, the old and new runs must remain available. Otherwise, memory turns a benchmark into folklore.
This is also where online claims fail. A press release may say a system has improved safety, while the relevant change was simply a wider operating domain or a new disengagement definition. The test question is narrower: did the vehicle make a better decision in the same situation, under the same constraints?
A practical workflow for engineers and owners
Start by downloading the original filing and preserving its publication date. Do not rely on a screenshot or a social-media summary. Build a local copy with the company, vehicle type, reporting period, miles, interventions, causes, and missing-data notes.
Next, calculate normalized rates and segment the exposure. Flag any comparison that mixes driverless service miles with supervised development miles. Then read individual narratives and assign severity only when the evidence supports it. Keep an uncertainty column for events that lack enough detail.
After that, compare the result with release notes, known operational boundaries, and your own repeatable drives. An owner can test a narrow question without pretending to validate an entire autonomy stack. For example, repeat the same protected-left-turn route before and after an update, using consistent weather and recording whether the car yields, creeps, or requests intervention.
Finally, publish the raw inputs beside the conclusion. A useful report lets another engineer reproduce the calculation and challenge the assumptions. If three runs produce three different answers, that is not an embarrassment. It may indicate traffic randomness, unstable localization, or a test that is not controlled enough.
What a useful conclusion sounds like
A credible av disengagement report analysis does not say that the company with the lowest count has the safest system. It says what was measured, what was not measured, how exposure was normalized, and which failure modes deserve reproduction.
My preferred conclusion has three parts: here is what happened, here is what the vehicle should have done, and here is the gap. That structure keeps the discussion attached to behavior instead of branding. It also makes room for an inconclusive result, which is often the honest result.
The next time a chart ranks an autonomous fleet, check the denominator, route mix, severity, and reporting definitions before sharing it. Then run the scenario yourself if you can. The spreadsheet is evidence. It is not the homework grade.