Forty-seven days. That's how long I waited between Tesla's FSD v12.4.3 and v12.5.1. Forty-seven days of running the old version through my test harness, logging its failures, documenting its hallucinations, and wondering: what will they actually fix?
On December 31, 2024, Tesla pushed v12.5.1 to my test vehicle. I ran the full battery of tests over the following week—same roads, same scenarios, same methodology. Same coffee, same cold Michigan mornings, same data logger recording every millisecond.
This is the first full OTA comparison on the site. Structured data. Before and after. What changed, what didn't, and what the numbers actually mean.
Before we dive in: if you haven't read the methodology post, go back and read it. I'm not going to re-explain the metrics here. This is the scorecard, not the rulebook.
Test Conditions
All tests conducted between December 15, 2024 and January 7, 2025, in and around Ann Arbor, Michigan.
Condition | v12.4.3 (Before) | v12.5.1 (After) |
|---|---|---|
Weather | Clear, 28–34°F | Clear, 22–30°F |
Road surface | Dry, salt-treated | Dry, salt-treated |
Time of day | 10:00–15:00 EST | 10:00–15:00 EST |
Route | Standardized 47-mile loop (highway, city, suburban, MCity perimeter) | Same |
Number of runs | 12 | 12 |
Test scenarios | 62 (full suite) | 62 (full suite) |
Important caveat: I run the same physical route and the same simulation scenarios, but traffic conditions vary. I can't control other drivers, pedestrians, or weather. I've tried to keep the test environment as consistent as possible, but these are real roads—not a lab. I note the conditions on each run and flag any outliers. For v12.5.1, I did three extra runs on a single day with near-identical traffic to validate the key metrics.
The Scorecard: Full Comparison
Metric | v12.4.3 (Before) | v12.5.1 (After) | Change | Verdict |
|---|---|---|---|---|
End-to-End Latency (95th percentile) | 147 ms | 112 ms | -35 ms (-24%) | ✅ Improved |
Physical Impossibility Hallucinations (per 1,000 miles) | 0.04 | 0.02 | -50% | ✅ Improved |
Semantic Inconsistency Hallucinations (per 1,000 miles) | 0.17 | 0.09 | -47% | ✅ Improved |
Probabilistic Outlier Hallucinations (per 1,000 miles) | 0.22 | 0.18 | -18% | ✅ Slight improvement |
Trajectory Reasonableness (Composite 0-100) | 78.4 | 83.1 | +4.7 | ✅ Improved |
— Curvature Continuity (30% weight) | 72.1 | 79.4 | +7.3 | ✅ Significant |
— Safety Margin (25% weight) | 84.2 | 85.6 | +1.4 | ↔️ Minimal |
— Action-Outcome Consistency (20% weight) | 76.8 | 82.0 | +5.2 | ✅ Improved |
— Human Likeness (15% weight) | 74.5 | 79.1 | +4.6 | ✅ Improved |
— Responsiveness (10% weight) | 85.3 | 86.7 | +1.4 | ↔️ Minimal |
Long-Tail Generalization (Composite % pass rate) | 76% | 84% | +8% | ✅ Improved |
— Adverse Weather & Lighting | 62% | 74% | +12% | ✅ Significant |
— Infrastructure Degradation | 71% | 78% | +7% | ✅ Improved |
— Agent Behavior Extremes | 79% | 83% | +4% | ↔️ Slight |
— Traffic Control Failures | 78% | 86% | +8% | ✅ Improved |
— Map/Odometer Discrepancies | 83% | 89% | +6% | ✅ Improved |
— "Weird" Scenarios | 67% | 72% | +5% | ↔️ Slight |
Tool Calling Accuracy (navigation commands) | 91.2% | 94.7% | +3.5% | ✅ Improved |
Tool Calling Accuracy (lane-change requests) | 88.3% | 92.1% | +3.8% | ✅ Improved |
What Changed: The Deep Dive
Latency: -35 ms
This is the headline number, and it's real. The v12.5.1 network is faster to produce trajectory outputs.
I don't have access to the model architecture, but my logs show a consistent reduction in inference time across all scenarios. My working hypothesis is that Tesla pruned the network—removed redundant layers or quantized weights—to reduce the compute cost at inference time. The latencies are now in the same ballpark as the best open-source driving models I've tested, which means Tesla may be approaching the hardware limits of the current-generation FSD computer.
The implication: A 112 ms 95th percentile latency means the car is processing sensor data and producing a trajectory in roughly the time it takes a human to perceive an event and begin a reaction. The industry standard for human driver perception-reaction time is around 150-200 ms for a visual stimulus. So in pure processing terms, the car is now faster than a human, but just barely. The gap used to be wider. Now it's a race.
Curvature Continuity: +7.3 points

This was the largest single improvement, and it's visible in the driving feel. The v12.5.1 trajectories have significantly lower jerk—the rate of change of acceleration.
In plain English: the car is smoother.
My CAN logs show that the steering commands are less aggressive in the v12.5.1 runs. The maximum steering torque peaks are lower (0.8 N·m vs. 1.4 N·m on the same section of I-94), and the changes in torque are gentler.
Why does this matter? Because smoothness isn't just about comfort. It's about communication. A driver who sees a car with smooth, progressive steering movements can predict its intent better than a car that suddenly jerks the wheel. The model is now communicating its decisions more clearly.
Adverse Weather & Lighting: +12%
This is the biggest improvement in the long-tail scenario suite, and it's the one I'm most skeptical about.
The v12.5.1 model handled my dusk low-light tests significantly better than the v12.4.3 version. It misclassified fewer pedestrians, stayed more confidently within lane lines, and didn't hallucinate merge lanes in shadow conditions.
The most dramatic improvement came on a dusk test I ran on US-23 at 6:15 PM, with the sun at the driver's 10 o'clock. On v12.4.3, the car had downgraded a pedestrian classification at 88 meters—the same condition that triggered Hallucination #102. On v12.5.1, the model maintained pedestrian classification at 92% confidence throughout the sequence.
I don't know if Tesla retrained the network on a larger dataset with low-light conditions, or if they changed the camera exposure parameters, or if they added a separate low-light model. The logs show a difference; I don't know the cause.
What hasn't changed (and why it matters):
The Safety Margin metric improved by a negligible 1.4 points. And the Probabilistic Outlier hallucination rate dropped only 18%.
What this tells me is that Tesla is focusing on the reproducible failures—the edge cases they can recreate in simulation and retrain on. The semantic inconsistencies and physical impossibilities are getting cleaned up. But the weird, rare, one-off behaviors—the model swerving 18 inches for no reason, the 3-sigma trajectory distributions—are harder to fix.
Those require architectural changes, not just dataset expansions.
What Didn't Change (The Plot Thickens)
The Safety Margin metric showed almost no improvement. The v12.5.1 car keeps the same minimum distance to objects and lane boundaries as the v12.4.3.
This tells me something important: Tesla's core safety envelope hasn't changed. The model's safety policy—the distance it keeps from obstacles, the sensitivity of the safety monitor, the trade-offs it makes—is fundamentally the same in v12.5.1.
A 1.4-point improvement in Safety Margin is noise. The car is keeping the same physical margins to the world.
What changed is the how—smoother trajectories, faster reactions, better semantic understanding—but the safety envelope stays fixed.
This is a wise engineering decision. Changing a safety envelope is risky. You don't tweak the minimum distance to obstacles in an OTA unless you're absolutely sure about the consequences. The model's safety policy is embedded in a combination of explicit constraints (e.g., "don't cross this boundary") and learned heuristics.
FSD v12.5.1 is a refinement, not a rewrite. The model's understanding of the road, the physics, and the semantics of the scene has improved. The safety policy hasn't changed, which means the car is making smarter decisions within the same safety constraints.
The Discrepancy that Bothers Me
Here's a mystery.
The Tool Calling Accuracy metric improved by 3-4 percentage points. The model is better at understanding lane-change requests from the driver. It's better at navigating.
But in the "Weird" scenarios category, the improvement was only 5 percentage points.
Why would the model improve at understanding driver commands but struggle with unusual road situations? The answer is likely in the training data distribution. Tesla is training on a massive dataset of real-world driving. That dataset is dominated by normal driving scenarios—highways, intersections, lane changes, stops. It's not dominated by weird scenarios—deer crossings, mattress-in-the-road, farm equipment.
The model's performance improves where the dataset is dense. It struggles where the dataset is sparse.
The driver commands are a known domain. The weird scenarios are not.