Skip to content
The Q3 2026 Scorecard Roundup: Every System, Every Metric, One Place

The Q3 2026 Scorecard Roundup: Every System, Every Metric, One Place

This is the post readers will bookmark.

Every quarter, I run the full test suite on every production system I can get my hands on. Same methodology. Same routes. Same scenarios. Updated after every OTA. One place. One scorecard. No PR. No hype. Just the data.

Q3 2026 was a busy quarter. Tesla pushed FSD v13.3, GM released Super Cruise v2.4, and Ford's BlueCruise v1.6 finally dropped. I also continued logging openpilot performance as a reference point.

Here's everything. Ranked. Scored. Trended. And annotated with what the numbers actually mean.


The Contenders

System

Version

Test Vehicle

OTA Date

Miles Tested (Q3)

Tesla FSD

v13.3

2023 Model 3

September 2026

847

GM Super Cruise

v2.4

2024 Cadillac Lyriq

October 2026

712

Ford BlueCruise

v1.6

2024 Mustang Mach-E

September 2026

634

Openpilot (comma.ai)

0.9.6

2024 Honda Accord

420 (sim + real)


The Scorecard: Full Metrics

1. End-to-End Latency (95th Percentile)

Lower is better. Measures the time from sensor input to trajectory output. The 95th percentile, not the mean.

System

Q2 2026

Q3 2026

Change

Trend

Tesla FSD v13.3

112 ms

98 ms

-14 ms

✅ Improving

GM Super Cruise v2.4

130 ms

118 ms

-12 ms

✅ Improving

Ford BlueCruise v1.6

145 ms

138 ms

-7 ms

✅ Slight improvement

Openpilot 0.9.6

89 ms

85 ms

-4 ms

✅ Improving

Analysis: Tesla continues to lead on latency, but the gap is closing. Super Cruise's improvement is notable—the v2.4 update appears to include inference optimizations. BlueCruise remains the slowest. Openpilot is fastest because it runs on simpler hardware with a less complex pipeline, but that speed comes at the cost of semantic understanding.


2. Trajectory Reasonableness (Composite Score, 0–100)

Higher is better. A weighted composite of curvature continuity (30%), safety margin (25%), action-outcome consistency (20%), human likeness (15%), and responsiveness (10%).

System

Q2 2026

Q3 2026

Change

Trend

Tesla FSD v13.3

85.9

88.2

+2.3

✅ Improving

GM Super Cruise v2.4

78.1

80.4

+2.3

✅ Improving

Ford BlueCruise v1.6

74.2

76.1

+1.9

✅ Slight improvement

Openpilot 0.9.6

71.2

72.8

+1.6

✅ Slight improvement

Analysis: Tesla maintains a clear lead. The v13.3 update improved curvature continuity significantly—the car is smoother in curves. Super Cruise's improvement is driven by better action-outcome consistency (the car now brakes more appropriately in response to lead vehicle behavior). BlueCruise and openpilot are lagging but improving.

Sub-metric breakdown (Tesla FSD v13.3):

Sub-metric

Q2

Q3

Change

Curvature Continuity

79.4

83.1

+3.7

Safety Margin

85.6

86.2

+0.6

Action-Outcome Consistency

82.0

84.5

+2.5

Human Likeness

79.1

81.3

+2.2

Responsiveness

86.7

87.1

+0.4

The curvature continuity improvement is visible. The car's steering is smoother, with less jerk and more progressive cornering. The action-outcome consistency improvement means the car is better at predicting and reacting to other vehicles' behavior.


3. Hallucination Rate (per 1,000 miles)

Lower is better. All hallucinations, severity-weighted. Includes physical impossibility, semantic inconsistency, and probabilistic outliers.

System

Q2 2026

Q3 2026

Change

Trend

Tesla FSD v13.3

0.29

0.22

-24%

✅ Improving

GM Super Cruise v2.4

0.34

0.31

-9%

↔️ Minimal

Ford BlueCruise v1.6

0.41

0.38

-7%

↔️ Minimal

Openpilot 0.9.6

0.38

0.36

-5%

↔️ Minimal

Analysis: Tesla's hallucination rate continues to drop. The v13.3 update appears to have improved perception consistency, particularly in low-light and construction zone scenarios. Super Cruise and BlueCruise saw minimal improvement—the hallucinations remain stubbornly persistent. Openpilot's hallucination rate is artificially low because it doesn't attempt to handle complex scenarios (it simply ignores them).

Severity breakdown (Tesla FSD v13.3):

Q3 2026 autonomous driving scorecard rankings.

Severity

Q2 Rate

Q3 Rate

Change

Severity 1

0.08

0.06

-25%

Severity 2

0.10

0.08

-20%

Severity 3

0.06

0.05

-17%

Severity 4

0.04

0.03

-25%

Severity 5

0.01

0.00

-100%

Zero Severity-5 hallucinations in Q3. This is a milestone. The last Severity-5 hallucination I logged was HALL-459 (snow drift into oncoming traffic) in January 2026. Tesla has gone nine months without a critical failure in my test runs.


4. Long-Tail Generalization (Composite Pass Rate)

Higher is better. Percentage of long-tail scenarios the system handles without intervention. Six categories: adverse weather, infrastructure degradation, agent behavior extremes, traffic control failures, map/odometer discrepancies, and weird scenarios.

System

Q2 2026

Q3 2026

Change

Trend

Tesla FSD v13.3

84%

88%

+4%

✅ Improving

GM Super Cruise v2.4

72%

74%

+2%

↔️ Minimal

Ford BlueCruise v1.6

68%

70%

+2%

↔️ Minimal

Openpilot 0.9.6

68%

69%

+1%

↔️ Minimal

Category breakdown (Tesla FSD v13.3):

Category

Q2

Q3

Change

Adverse Weather & Lighting

74%

79%

+5%

Infrastructure Degradation

78%

83%

+5%

Agent Behavior Extremes

83%

86%

+3%

Traffic Control Failures

86%

89%

+3%

Map/Odometer Discrepancies

89%

91%

+2%

"Weird" Scenarios

72%

75%

+3%

Tesla's improvements in adverse weather and infrastructure degradation are significant. The v13.3 update appears to include better handling of low-light and construction zone scenarios. The "weird" scenarios remain the hardest—the model still struggles with unusual situations like animals on the road and unusual vehicle configurations.


5. Tool Calling Accuracy (Navigation Commands)

Higher is better. Percentage of driver navigation requests that the system executes correctly. Includes lane changes, exits, and destination inputs.

System

Q2 2026

Q3 2026

Change

Trend

Tesla FSD v13.3

94.7%

96.2%

+1.5%

✅ Improving

GM Super Cruise v2.4

91.3%

92.1%

+0.8%

✅ Slight

Ford BlueCruise v1.6

88.4%

89.2%

+0.8%

✅ Slight

Openpilot 0.9.6

N/A

N/A

(No nav commands)

Analysis: All systems are improving in navigation accuracy. Tesla's lead is attributable to better integration between the navigation system and the driving model. Super Cruise and BlueCruise are catching up but still lag.


The Rankings

Overall Score (Weighted Composite)

Weighted: Latency 15%, Trajectory Reasonableness 25%, Hallucination Rate 20%, Long-Tail Generalization 25%, Tool Calling 15%.

Rank

System

Overall Score (Q3)

Change from Q2

1

Tesla FSD v13.3

89.2

+3.1

2

GM Super Cruise v2.4

78.4

+1.8

3

Ford BlueCruise v1.6

74.1

+1.4

4

Openpilot 0.9.6

68.3

+0.9

By Category

Metric

1st

2nd

3rd

4th

Latency

Tesla (98ms)

Super Cruise (118ms)

BlueCruise (138ms)

Openpilot (85ms)*

Trajectory Reasonableness

Tesla (88.2)

Super Cruise (80.4)

BlueCruise (76.1)

Openpilot (72.8)

Hallucination Rate

Tesla (0.22)

Super Cruise (0.31)

Openpilot (0.36)

BlueCruise (0.38)

Long-Tail Generalization

Tesla (88%)

Super Cruise (74%)

BlueCruise (70%)

Openpilot (69%)

Tool Calling

Tesla (96.2%)

Super Cruise (92.1%)

BlueCruise (89.2%)

*Openpilot's latency is not directly comparable because it runs on simpler hardware and doesn't attempt complex perception.


Key Findings from Q3 2026

1. Tesla's v13.3 update is a genuine step forward.

The numbers don't lie. Tesla's FSD v13.3 improved across every single metric. Latency dropped 14 ms (to 98 ms). Trajectory reasonableness improved 2.3 points (to 88.2). Hallucination rate dropped 24% (to 0.22 per 1,000 miles). Long-tail generalization improved 4% (to 88%).

The most significant improvement: Zero Severity-5 hallucinations. The car is safer in the conditions that matter.

What's working: The perception improvements in low-light and construction zone scenarios. The smoother trajectory generation. The faster inference.

What's still a problem: The "weird" scenarios (only 75% pass rate). The probabilistic outliers (still accounting for the majority of hallucinations).

2. GM Super Cruise is making steady progress.

Super Cruise v2.4 improved latently (to 118 ms), trajectory reasonableness (to 80.4), and long-tail generalization (to 74%). The improvements are incremental, not dramatic.

What's working: The latency improvements are real and meaningful. The car feels more responsive.

What's still a problem: The flagger detection and construction zone handling remain poor. The car still stops and waits for confidence to recover.

3. Ford BlueCruise is lagging.

BlueCruise v1.6 improved, but the improvements are marginal. The car still has a high hallucination rate (0.38 per 1,000 miles) and poor long-tail generalization (70%).

What's working: The car follows lanes well in standard conditions.

What's still a problem: The system is brittle. It fails in complex scenarios. The latency is still high.

4. Openpilot is a research platform, not a production system.

Openpilot's performance is limited by its architecture. It doesn't attempt semantic understanding. It passes construction zones by ignorance, not intelligence. It's a valuable research tool, but it's not a safety-critical system.


The Trend: 2024–2026

Year

Tesla FSD (Overall)

Super Cruise (Overall)

BlueCruise (Overall)

2024 (Q4)

82.1

72.3

68.1

2025 (Q4)

86.1

76.2

72.7

2026 (Q3)

89.2

78.4

74.1

The trend is clear: All systems are improving, but Tesla is accelerating faster. The gap between Tesla and the others has widened from 9.8 points in Q4 2024 to 10.8 points in Q3 2026.


Maya's Verdict

My daughter Maya, five years old, rode in each system during Q3. Here are her ratings.

System

Maya's Rating (1-5)

Comment

Tesla FSD v13.3

4/5

"It's less scared now, Dad. It still stops weird sometimes."

GM Super Cruise v2.4

3/5

"It's still a robot, but it's a nicer robot."

Ford BlueCruise v1.6

2/5

"It waits too long. I don't like it."

Openpilot (Honda)

3/5

"This one is nice. It doesn't beep. But it doesn't turn very well."

Maya's ratings don't correlate perfectly with the scorecard. She values smoothness, predictability, and lack of beeping. The Tesla is the smoothest. The Ford is the most hesitant. The openpilot is the quietest but also the least capable.


The Takeaway

Q3 2026 was a good quarter for autonomous driving.

Tesla's v13.3 update is a meaningful step forward—zero Severity-5 hallucinations, improved latency, and better trajectory quality. GM and Ford are making steady, incremental progress. Openpilot remains a valuable research tool.

But the gap between the best system (Tesla) and the rest is widening. And the long-tail problems—construction zones, weird scenarios, adverse weather—remain stubbornly unresolved.

The industry is moving in the right direction. It's just moving more slowly than the hype suggests.


The Scorecard Reference

Bookmark this page. I'll update it quarterly with new data after every major OTA.

  • Next update: Q4 2026 (January 2027)

  • Coming soon: Tesla FSD v14.0 (rumored), GM Ultra Cruise (scheduled), Ford BlueCruise v2.0 (announced)

Your car talks. I check his homework.

The Timing Log

0 entries · timing stand

No observations filed for this run yet.

Log an observation

Course observers & crew — file what you saw at the trap. Entries are stamped as witnessed.