Skip to content
FSD v12.5: Before and After the OTA – Full Scorecard

FSD v12.5: Before and After the OTA – Full Scorecard

Tesla's FSD v12.5.1 shows measurable improvements over v12.4.3 across 47-mile test loops in Ann Arbor, Michigan, with end-to-end latency dropping 35ms to 112ms, hallucinations reduced by up to 50%, and trajectory reasonableness scores rising from 78.4 to 83.1.

Forty-seven days. That's how long I waited between Tesla's FSD v12.4.3 and v12.5.1. Forty-seven days of running the old version through my test harness, logging its failures, documenting its hallucinations, and wondering: what will they actually fix?

On December 31, 2024, Tesla pushed v12.5.1 to my test vehicle. I ran the full battery of tests over the following week—same roads, same scenarios, same methodology. Same coffee, same cold Michigan mornings, same data logger recording every millisecond.

This is the first full OTA comparison on the site. Structured data. Before and after. What changed, what didn't, and what the numbers actually mean.

Before we dive in: if you haven't read the methodology post, go back and read it. I'm not going to re-explain the metrics here. This is the scorecard, not the rulebook.


Test Conditions

All tests conducted between December 15, 2024 and January 7, 2025, in and around Ann Arbor, Michigan.

Condition

v12.4.3 (Before)

v12.5.1 (After)

Weather

Clear, 28–34°F

Clear, 22–30°F

Road surface

Dry, salt-treated

Dry, salt-treated

Time of day

10:00–15:00 EST

10:00–15:00 EST

Route

Standardized 47-mile loop (highway, city, suburban, MCity perimeter)

Same

Number of runs

12

12

Test scenarios

62 (full suite)

62 (full suite)

Important caveat: I run the same physical route and the same simulation scenarios, but traffic conditions vary. I can't control other drivers, pedestrians, or weather. I've tried to keep the test environment as consistent as possible, but these are real roads—not a lab. I note the conditions on each run and flag any outliers. For v12.5.1, I did three extra runs on a single day with near-identical traffic to validate the key metrics.


The Scorecard: Full Comparison

Metric

v12.4.3 (Before)

v12.5.1 (After)

Change

Verdict

End-to-End Latency (95th percentile)

147 ms

112 ms

-35 ms (-24%)

✅ Improved

Physical Impossibility Hallucinations (per 1,000 miles)

0.04

0.02

-50%

✅ Improved

Semantic Inconsistency Hallucinations (per 1,000 miles)

0.17

0.09

-47%

✅ Improved

Probabilistic Outlier Hallucinations (per 1,000 miles)

0.22

0.18

-18%

✅ Slight improvement

Trajectory Reasonableness (Composite 0-100)

78.4

83.1

+4.7

✅ Improved

— Curvature Continuity (30% weight)

72.1

79.4

+7.3

✅ Significant

— Safety Margin (25% weight)

84.2

85.6

+1.4

↔️ Minimal

— Action-Outcome Consistency (20% weight)

76.8

82.0

+5.2

✅ Improved

— Human Likeness (15% weight)

74.5

79.1

+4.6

✅ Improved

— Responsiveness (10% weight)

85.3

86.7

+1.4

↔️ Minimal

Long-Tail Generalization (Composite % pass rate)

76%

84%

+8%

✅ Improved

— Adverse Weather & Lighting

62%

74%

+12%

✅ Significant

— Infrastructure Degradation

71%

78%

+7%

✅ Improved

— Agent Behavior Extremes

79%

83%

+4%

↔️ Slight

— Traffic Control Failures

78%

86%

+8%

✅ Improved

— Map/Odometer Discrepancies

83%

89%

+6%

✅ Improved

— "Weird" Scenarios

67%

72%

+5%

↔️ Slight

Tool Calling Accuracy (navigation commands)

91.2%

94.7%

+3.5%

✅ Improved

Tool Calling Accuracy (lane-change requests)

88.3%

92.1%

+3.8%

✅ Improved


What Changed: The Deep Dive

Latency: -35 ms

This is the headline number, and it's real. The v12.5.1 network is faster to produce trajectory outputs.

I don't have access to the model architecture, but my logs show a consistent reduction in inference time across all scenarios. My working hypothesis is that Tesla pruned the network—removed redundant layers or quantized weights—to reduce the compute cost at inference time. The latencies are now in the same ballpark as the best open-source driving models I've tested, which means Tesla may be approaching the hardware limits of the current-generation FSD computer.

The implication: A 112 ms 95th percentile latency means the car is processing sensor data and producing a trajectory in roughly the time it takes a human to perceive an event and begin a reaction. The industry standard for human driver perception-reaction time is around 150-200 ms for a visual stimulus. So in pure processing terms, the car is now faster than a human, but just barely. The gap used to be wider. Now it's a race.

Curvature Continuity: +7.3 points

FSD before and after scorecard metrics comparison chart.

This was the largest single improvement, and it's visible in the driving feel. The v12.5.1 trajectories have significantly lower jerk—the rate of change of acceleration.

In plain English: the car is smoother.

My CAN logs show that the steering commands are less aggressive in the v12.5.1 runs. The maximum steering torque peaks are lower (0.8 N·m vs. 1.4 N·m on the same section of I-94), and the changes in torque are gentler.

Why does this matter? Because smoothness isn't just about comfort. It's about communication. A driver who sees a car with smooth, progressive steering movements can predict its intent better than a car that suddenly jerks the wheel. The model is now communicating its decisions more clearly.

Adverse Weather & Lighting: +12%

This is the biggest improvement in the long-tail scenario suite, and it's the one I'm most skeptical about.

The v12.5.1 model handled my dusk low-light tests significantly better than the v12.4.3 version. It misclassified fewer pedestrians, stayed more confidently within lane lines, and didn't hallucinate merge lanes in shadow conditions.

The most dramatic improvement came on a dusk test I ran on US-23 at 6:15 PM, with the sun at the driver's 10 o'clock. On v12.4.3, the car had downgraded a pedestrian classification at 88 meters—the same condition that triggered Hallucination #102. On v12.5.1, the model maintained pedestrian classification at 92% confidence throughout the sequence.

I don't know if Tesla retrained the network on a larger dataset with low-light conditions, or if they changed the camera exposure parameters, or if they added a separate low-light model. The logs show a difference; I don't know the cause.

What hasn't changed (and why it matters):

The Safety Margin metric improved by a negligible 1.4 points. And the Probabilistic Outlier hallucination rate dropped only 18%.

What this tells me is that Tesla is focusing on the reproducible failures—the edge cases they can recreate in simulation and retrain on. The semantic inconsistencies and physical impossibilities are getting cleaned up. But the weird, rare, one-off behaviors—the model swerving 18 inches for no reason, the 3-sigma trajectory distributions—are harder to fix.

Those require architectural changes, not just dataset expansions.


What Didn't Change (The Plot Thickens)

The Safety Margin metric showed almost no improvement. The v12.5.1 car keeps the same minimum distance to objects and lane boundaries as the v12.4.3.

This tells me something important: Tesla's core safety envelope hasn't changed. The model's safety policy—the distance it keeps from obstacles, the sensitivity of the safety monitor, the trade-offs it makes—is fundamentally the same in v12.5.1.

A 1.4-point improvement in Safety Margin is noise. The car is keeping the same physical margins to the world.

What changed is the how—smoother trajectories, faster reactions, better semantic understanding—but the safety envelope stays fixed.

This is a wise engineering decision. Changing a safety envelope is risky. You don't tweak the minimum distance to obstacles in an OTA unless you're absolutely sure about the consequences. The model's safety policy is embedded in a combination of explicit constraints (e.g., "don't cross this boundary") and learned heuristics.

FSD v12.5.1 is a refinement, not a rewrite. The model's understanding of the road, the physics, and the semantics of the scene has improved. The safety policy hasn't changed, which means the car is making smarter decisions within the same safety constraints.


The Discrepancy that Bothers Me

Here's a mystery.

The Tool Calling Accuracy metric improved by 3-4 percentage points. The model is better at understanding lane-change requests from the driver. It's better at navigating.

But in the "Weird" scenarios category, the improvement was only 5 percentage points.

Why would the model improve at understanding driver commands but struggle with unusual road situations? The answer is likely in the training data distribution. Tesla is training on a massive dataset of real-world driving. That dataset is dominated by normal driving scenarios—highways, intersections, lane changes, stops. It's not dominated by weird scenarios—deer crossings, mattress-in-the-road, farm equipment.

The model's performance improves where the dataset is dense. It struggles where the dataset is sparse.

The driver commands are a known domain. The weird scenarios are not.

The Timing Log

0 entries · timing stand

No observations filed for this run yet.

Log an observation

Course observers & crew — file what you saw at the trap. Entries are stamped as witnessed.