This is the post readers will bookmark.
Every quarter, I run the full test suite on every production system I can get my hands on. Same methodology. Same routes. Same scenarios. Updated after every OTA. One place. One scorecard. No PR. No hype. Just the data.
Q3 2026 was a busy quarter. Tesla pushed FSD v13.3, GM released Super Cruise v2.4, and Ford's BlueCruise v1.6 finally dropped. I also continued logging openpilot performance as a reference point.
Here's everything. Ranked. Scored. Trended. And annotated with what the numbers actually mean.
The Contenders
System | Version | Test Vehicle | OTA Date | Miles Tested (Q3) |
|---|---|---|---|---|
Tesla FSD | v13.3 | 2023 Model 3 | September 2026 | 847 |
GM Super Cruise | v2.4 | 2024 Cadillac Lyriq | October 2026 | 712 |
Ford BlueCruise | v1.6 | 2024 Mustang Mach-E | September 2026 | 634 |
Openpilot (comma.ai) | 0.9.6 | 2024 Honda Accord | — | 420 (sim + real) |
The Scorecard: Full Metrics
1. End-to-End Latency (95th Percentile)
Lower is better. Measures the time from sensor input to trajectory output. The 95th percentile, not the mean.
System | Q2 2026 | Q3 2026 | Change | Trend |
|---|---|---|---|---|
Tesla FSD v13.3 | 112 ms | 98 ms | -14 ms | ✅ Improving |
GM Super Cruise v2.4 | 130 ms | 118 ms | -12 ms | ✅ Improving |
Ford BlueCruise v1.6 | 145 ms | 138 ms | -7 ms | ✅ Slight improvement |
Openpilot 0.9.6 | 89 ms | 85 ms | -4 ms | ✅ Improving |
Analysis: Tesla continues to lead on latency, but the gap is closing. Super Cruise's improvement is notable—the v2.4 update appears to include inference optimizations. BlueCruise remains the slowest. Openpilot is fastest because it runs on simpler hardware with a less complex pipeline, but that speed comes at the cost of semantic understanding.
2. Trajectory Reasonableness (Composite Score, 0–100)
Higher is better. A weighted composite of curvature continuity (30%), safety margin (25%), action-outcome consistency (20%), human likeness (15%), and responsiveness (10%).
System | Q2 2026 | Q3 2026 | Change | Trend |
|---|---|---|---|---|
Tesla FSD v13.3 | 85.9 | 88.2 | +2.3 | ✅ Improving |
GM Super Cruise v2.4 | 78.1 | 80.4 | +2.3 | ✅ Improving |
Ford BlueCruise v1.6 | 74.2 | 76.1 | +1.9 | ✅ Slight improvement |
Openpilot 0.9.6 | 71.2 | 72.8 | +1.6 | ✅ Slight improvement |
Analysis: Tesla maintains a clear lead. The v13.3 update improved curvature continuity significantly—the car is smoother in curves. Super Cruise's improvement is driven by better action-outcome consistency (the car now brakes more appropriately in response to lead vehicle behavior). BlueCruise and openpilot are lagging but improving.
Sub-metric breakdown (Tesla FSD v13.3):
Sub-metric | Q2 | Q3 | Change |
|---|---|---|---|
Curvature Continuity | 79.4 | 83.1 | +3.7 |
Safety Margin | 85.6 | 86.2 | +0.6 |
Action-Outcome Consistency | 82.0 | 84.5 | +2.5 |
Human Likeness | 79.1 | 81.3 | +2.2 |
Responsiveness | 86.7 | 87.1 | +0.4 |
The curvature continuity improvement is visible. The car's steering is smoother, with less jerk and more progressive cornering. The action-outcome consistency improvement means the car is better at predicting and reacting to other vehicles' behavior.
3. Hallucination Rate (per 1,000 miles)
Lower is better. All hallucinations, severity-weighted. Includes physical impossibility, semantic inconsistency, and probabilistic outliers.
System | Q2 2026 | Q3 2026 | Change | Trend |
|---|---|---|---|---|
Tesla FSD v13.3 | 0.29 | 0.22 | -24% | ✅ Improving |
GM Super Cruise v2.4 | 0.34 | 0.31 | -9% | ↔️ Minimal |
Ford BlueCruise v1.6 | 0.41 | 0.38 | -7% | ↔️ Minimal |
Openpilot 0.9.6 | 0.38 | 0.36 | -5% | ↔️ Minimal |
Analysis: Tesla's hallucination rate continues to drop. The v13.3 update appears to have improved perception consistency, particularly in low-light and construction zone scenarios. Super Cruise and BlueCruise saw minimal improvement—the hallucinations remain stubbornly persistent. Openpilot's hallucination rate is artificially low because it doesn't attempt to handle complex scenarios (it simply ignores them).
Severity breakdown (Tesla FSD v13.3):

Severity | Q2 Rate | Q3 Rate | Change |
|---|---|---|---|
Severity 1 | 0.08 | 0.06 | -25% |
Severity 2 | 0.10 | 0.08 | -20% |
Severity 3 | 0.06 | 0.05 | -17% |
Severity 4 | 0.04 | 0.03 | -25% |
Severity 5 | 0.01 | 0.00 | -100% |
Zero Severity-5 hallucinations in Q3. This is a milestone. The last Severity-5 hallucination I logged was HALL-459 (snow drift into oncoming traffic) in January 2026. Tesla has gone nine months without a critical failure in my test runs.
4. Long-Tail Generalization (Composite Pass Rate)
Higher is better. Percentage of long-tail scenarios the system handles without intervention. Six categories: adverse weather, infrastructure degradation, agent behavior extremes, traffic control failures, map/odometer discrepancies, and weird scenarios.
System | Q2 2026 | Q3 2026 | Change | Trend |
|---|---|---|---|---|
Tesla FSD v13.3 | 84% | 88% | +4% | ✅ Improving |
GM Super Cruise v2.4 | 72% | 74% | +2% | ↔️ Minimal |
Ford BlueCruise v1.6 | 68% | 70% | +2% | ↔️ Minimal |
Openpilot 0.9.6 | 68% | 69% | +1% | ↔️ Minimal |
Category breakdown (Tesla FSD v13.3):
Category | Q2 | Q3 | Change |
|---|---|---|---|
Adverse Weather & Lighting | 74% | 79% | +5% |
Infrastructure Degradation | 78% | 83% | +5% |
Agent Behavior Extremes | 83% | 86% | +3% |
Traffic Control Failures | 86% | 89% | +3% |
Map/Odometer Discrepancies | 89% | 91% | +2% |
"Weird" Scenarios | 72% | 75% | +3% |
Tesla's improvements in adverse weather and infrastructure degradation are significant. The v13.3 update appears to include better handling of low-light and construction zone scenarios. The "weird" scenarios remain the hardest—the model still struggles with unusual situations like animals on the road and unusual vehicle configurations.
5. Tool Calling Accuracy (Navigation Commands)
Higher is better. Percentage of driver navigation requests that the system executes correctly. Includes lane changes, exits, and destination inputs.
System | Q2 2026 | Q3 2026 | Change | Trend |
|---|---|---|---|---|
Tesla FSD v13.3 | 94.7% | 96.2% | +1.5% | ✅ Improving |
GM Super Cruise v2.4 | 91.3% | 92.1% | +0.8% | ✅ Slight |
Ford BlueCruise v1.6 | 88.4% | 89.2% | +0.8% | ✅ Slight |
Openpilot 0.9.6 | N/A | N/A | — | (No nav commands) |
Analysis: All systems are improving in navigation accuracy. Tesla's lead is attributable to better integration between the navigation system and the driving model. Super Cruise and BlueCruise are catching up but still lag.
The Rankings
Overall Score (Weighted Composite)
Weighted: Latency 15%, Trajectory Reasonableness 25%, Hallucination Rate 20%, Long-Tail Generalization 25%, Tool Calling 15%.
Rank | System | Overall Score (Q3) | Change from Q2 |
|---|---|---|---|
1 | Tesla FSD v13.3 | 89.2 | +3.1 |
2 | GM Super Cruise v2.4 | 78.4 | +1.8 |
3 | Ford BlueCruise v1.6 | 74.1 | +1.4 |
4 | Openpilot 0.9.6 | 68.3 | +0.9 |
By Category
Metric | 1st | 2nd | 3rd | 4th |
|---|---|---|---|---|
Latency | Tesla (98ms) | Super Cruise (118ms) | BlueCruise (138ms) | Openpilot (85ms)* |
Trajectory Reasonableness | Tesla (88.2) | Super Cruise (80.4) | BlueCruise (76.1) | Openpilot (72.8) |
Hallucination Rate | Tesla (0.22) | Super Cruise (0.31) | Openpilot (0.36) | BlueCruise (0.38) |
Long-Tail Generalization | Tesla (88%) | Super Cruise (74%) | BlueCruise (70%) | Openpilot (69%) |
Tool Calling | Tesla (96.2%) | Super Cruise (92.1%) | BlueCruise (89.2%) | — |
*Openpilot's latency is not directly comparable because it runs on simpler hardware and doesn't attempt complex perception.
Key Findings from Q3 2026
1. Tesla's v13.3 update is a genuine step forward.
The numbers don't lie. Tesla's FSD v13.3 improved across every single metric. Latency dropped 14 ms (to 98 ms). Trajectory reasonableness improved 2.3 points (to 88.2). Hallucination rate dropped 24% (to 0.22 per 1,000 miles). Long-tail generalization improved 4% (to 88%).
The most significant improvement: Zero Severity-5 hallucinations. The car is safer in the conditions that matter.
What's working: The perception improvements in low-light and construction zone scenarios. The smoother trajectory generation. The faster inference.
What's still a problem: The "weird" scenarios (only 75% pass rate). The probabilistic outliers (still accounting for the majority of hallucinations).
2. GM Super Cruise is making steady progress.
Super Cruise v2.4 improved latently (to 118 ms), trajectory reasonableness (to 80.4), and long-tail generalization (to 74%). The improvements are incremental, not dramatic.
What's working: The latency improvements are real and meaningful. The car feels more responsive.
What's still a problem: The flagger detection and construction zone handling remain poor. The car still stops and waits for confidence to recover.
3. Ford BlueCruise is lagging.
BlueCruise v1.6 improved, but the improvements are marginal. The car still has a high hallucination rate (0.38 per 1,000 miles) and poor long-tail generalization (70%).
What's working: The car follows lanes well in standard conditions.
What's still a problem: The system is brittle. It fails in complex scenarios. The latency is still high.
4. Openpilot is a research platform, not a production system.
Openpilot's performance is limited by its architecture. It doesn't attempt semantic understanding. It passes construction zones by ignorance, not intelligence. It's a valuable research tool, but it's not a safety-critical system.
The Trend: 2024–2026
Year | Tesla FSD (Overall) | Super Cruise (Overall) | BlueCruise (Overall) |
|---|---|---|---|
2024 (Q4) | 82.1 | 72.3 | 68.1 |
2025 (Q4) | 86.1 | 76.2 | 72.7 |
2026 (Q3) | 89.2 | 78.4 | 74.1 |
The trend is clear: All systems are improving, but Tesla is accelerating faster. The gap between Tesla and the others has widened from 9.8 points in Q4 2024 to 10.8 points in Q3 2026.
Maya's Verdict
My daughter Maya, five years old, rode in each system during Q3. Here are her ratings.
System | Maya's Rating (1-5) | Comment |
|---|---|---|
Tesla FSD v13.3 | 4/5 | "It's less scared now, Dad. It still stops weird sometimes." |
GM Super Cruise v2.4 | 3/5 | "It's still a robot, but it's a nicer robot." |
Ford BlueCruise v1.6 | 2/5 | "It waits too long. I don't like it." |
Openpilot (Honda) | 3/5 | "This one is nice. It doesn't beep. But it doesn't turn very well." |
Maya's ratings don't correlate perfectly with the scorecard. She values smoothness, predictability, and lack of beeping. The Tesla is the smoothest. The Ford is the most hesitant. The openpilot is the quietest but also the least capable.
The Takeaway
Q3 2026 was a good quarter for autonomous driving.
Tesla's v13.3 update is a meaningful step forward—zero Severity-5 hallucinations, improved latency, and better trajectory quality. GM and Ford are making steady, incremental progress. Openpilot remains a valuable research tool.
But the gap between the best system (Tesla) and the rest is widening. And the long-tail problems—construction zones, weird scenarios, adverse weather—remain stubbornly unresolved.
The industry is moving in the right direction. It's just moving more slowly than the hype suggests.
The Scorecard Reference
Bookmark this page. I'll update it quarterly with new data after every major OTA.
Next update: Q4 2026 (January 2027)
Coming soon: Tesla FSD v14.0 (rumored), GM Ultra Cruise (scheduled), Ford BlueCruise v2.0 (announced)
Your car talks. I check his homework.