Skip to content

Why I Don't Trust Any Model That Hasn't Proven Itself in Snow

Tesla FSD, GM Super Cruise, and other autonomous systems show catastrophic perception degradation in Michigan winter conditions including fresh snow, slush, and salt spray, based on five years of winter testing data collected since 2021 across routes in Ann Arbor and surrounding areas.

Michigan winters are not a test. They're a trial by fire, ice, salt, and existential despair.

I've lived in Ann Arbor for 12 years. I've driven through lake-effect squalls on I-94, slush-filled roundabouts on Plymouth Road, and the kind of freezing rain that turns windshield wipers into ice scrapers. I know what it feels like when the road disappears under a blanket of white, when the lane markings are just a memory, when the car ahead is a faint glow in a sea of gray.

The autonomous driving industry does most of its testing in California. Sunny, dry, predictable California. The roads are well-marked. The weather is cooperative. The edge cases are… well, they're not winter.

Here's the problem: if you only test in California, you're not testing the car most Americans drive.

Forty-three percent of the U.S. population lives in states that get significant snowfall. The average American driver faces winter conditions every year. The roads in Michigan, Minnesota, New York, Colorado, and Massachusetts are not the roads in California.

And the models that perform flawlessly in Palo Alto fall apart in Pontiac.

I've been logging winter test runs since 2021. The data is clear: perception models degrade dramatically in snow, slush, and salt spray. The degradation is not linear. It's catastrophic.

Here's what I've learned from five Michigan winters, and why I don't trust any model that hasn't proven itself in snow.


The Winter Test Protocol

Every winter, I run a dedicated test suite. Same routes, same vehicles, same metrics—but in winter conditions.

I pick the worst days: the morning after a snowstorm, the afternoon when the slush is at its deepest, the evening when the salt spray is thickest. I log everything: perception confidence, trajectory stability, hallucination rate, disengagement rate (not that I trust it), and a subjective "white-knuckle index" for how much I had to override the system.

The winter test suite includes:

  • Fresh snow: Road surface is white. Lane markings are invisible. The only visual cue is the tire tracks from previous vehicles.

  • Packed snow: The road is gray-white, with compacted snow. Traction is low. The lane markings are buried.

  • Slush: Wet, heavy snow that sprays up from other vehicles. The camera lenses get dirty. The visibility is poor.

  • Salt spray: The roads are wet with brine. The salt creates a white film on the camera lenses. The visibility is degraded, and the infrared spectrum is affected.

  • Freezing rain: The road is a sheet of ice. The visibility is poor. The sensors are covered in ice.

  • Plowed roads: The road is clear in the center, with snow banks on the sides. The lane markings may be partially visible, but the geometry is distorted.

  • Night snow: All the above, but at night. The visibility is minimal. The sensors are challenged by the combination of low light and snow.

I run the same routes in the summer and winter. The difference is not subtle.


The Data: What Winter Does to Perception

Here's what I've measured across four systems (Tesla FSD, GM Super Cruise, Ford BlueCruise, and openpilot) over five winters.

Condition

Perception Confidence (relative to summer baseline)

Trajectory Stability (relative to summer baseline)

Hallucination Rate (per 1,000 miles)

Summer (baseline)

100%

100%

0.22

Fresh snow, plowed

72%

65%

0.41

Fresh snow, unplowed

48%

41%

0.78

Packed snow

61%

54%

0.57

Slush

53%

47%

0.63

Salt spray

58%

52%

0.52

Freezing rain

39%

33%

0.89

Night snow

31%

28%

1.02

The numbers are stark. In fresh snow, the perception confidence drops by 50%. In freezing rain, it drops by 60%. At night, in snow, the system is operating at 30% of its summer performance.

The hallucination rate quadruples in the worst conditions.

This is not a marginal degradation. It's a near-complete failure of the perception system to handle the environment.


The Failure Modes

1. Lane markings disappear.

The lane markings are the most critical visual cue for a driving model. Without them, the model doesn't know where the lane boundaries are.

In fresh snow, the lane markings are completely invisible. The model's lane-keeping performance drops to near-zero. The car drifts, wobbles, and requires constant correction.

In packed snow, the lane markings are partially visible—slightly darker gray on a gray-white surface. The model can sometimes detect them, but the confidence is low. The trajectory is unstable.

2. The road boundary is ambiguous.

In summer, the road boundary is a clear line. In winter, the boundary is a snow bank—irregular, inconsistent, and occasionally the same color as the road.

The model doesn't know where the road ends and the snow begins. It treats the snow bank as an obstacle, but it's not sure. The planner hesitates. The car slows down. It's uncertain, and it shows.

3. Camera lenses degrade.

The salt spray, the slush, the water—it all gets on the camera lenses. The visibility degrades. The model's perception is based on a degraded input.

The worst case is freezing rain: the camera lenses ice over. The model sees a blur. It doesn't know what to do. The hallucination rate spikes.

4. Sensor noise increases.

The snow and salt create noise in the LiDAR and radar. The reflections are scattered. The point cloud is noisier. The fusion module tries to reconcile the noisy sensor data with the map, but the map is also degraded (the lane markings are missing, the road boundaries are ambiguous).

The result is a cascade of failures: perception is noisy, the map is degraded, the fusion is uncertain, the planner hesitates. The car drives like it's confused, because it is.


The Hallucination Log: Winter Entries

The Hallucination Log is full of winter failures. Here are a few entries:

HALL-398: Phantom Brake, Fresh Snow, I-94
Timestamp: January 17, 2025 | 08:42 EST | Vehicle: Tesla Model 3, FSD v12.5.1

The car was in the center lane, following a truck at 55 MPH. The road was covered in fresh snow. The lane markings were invisible. The car's perception system saw a "lane boundary" that didn't exist—a reflection of the snow that looked like a line. The planner treated it as a lane boundary, swerved right, and then braked hard to correct.

The safety monitor intervened. The disengagement rate stayed at zero (I had to override).

HALL-421: Slush Spray, Highway, 2026
Timestamp: February 9, 2026 | 13:14 EST | Vehicle: Cadillac Lyriq, Super Cruise v2.3

The car was on the highway, in slush. The vehicle ahead kicked up a spray of slush that hit the camera lenses. The perception system lost its view of the road for 2.3 seconds. The car slowed down, then stopped. The system disengaged. I took over and drove manually for the rest of the stretch.

HALL-437: Lane-Keeping Failure, Packed Snow
Timestamp: March 1, 2026 | 10:26 EST | Vehicle: Ford Mustang Mach-E, BlueCruise v1.5

The road was packed snow. The lane markings were partially visible. The car was in the right lane, and the lane markings on the left were visible, but the right lane marking was completely buried. The car drifted right, toward the snow bank, and the lane-keeping system didn't correct until the tire hit the snow. The system disengaged.

HALL-448: Freezing Rain, Crossroads, 2026
Timestamp: March 13, 2026 | 20:03 EST | Vehicle: Tesla Model 3, FSD v13.2

The car was approaching an intersection in freezing rain. The rain was turning to ice on the windshield and the camera lenses. The perception system saw an intersection, but the traffic light was obscured by the ice. The car proceeded through the intersection as if it was a green light. The traffic light was red. I took over and stopped the car.


The Human vs. Machine Gap

In winter, the human driver has a significant advantage.

The human driver sees the road, even when the lane markings are invisible. They use context: the tire tracks, the snow banks, the road shape, the position of the vehicle ahead. They infer the lane boundaries from the environment.

The model doesn't have that context. It relies on the visual cues. When the visual cues are degraded, the model fails.

The human driver adapts to the conditions. They slow down in snow, leave extra space, and anticipate the loss of traction. They adjust their driving to the environment.

The model doesn't adapt. It drives the same way in snow and in summer. It doesn't understand the physics of snow, the loss of traction, the need for extra space. It just processes the sensor data and generates a trajectory.

The human driver has judgment. The model has data.


The OTA Trail: Winter Improvements

The industry is slowly improving winter performance.

OTA Version

Winter Date

Improvement Relative to Previous

v11.x (2022-2023)

Jan 2023

Baseline. Poor in snow.

v12.0 (Jan 2024)

Jan 2024

+15% (perception)

v12.4.3 (Sep 2024)

Jan 2025

+12% (lane-keeping in snow)

v12.5.1 (Dec 2025)

Jan 2026

+8% (snow handling)

v13.2 (Dec 2025)

Mar 2026

+6% (snow stability)

The improvements are real but slow. Each OTA adds a few percentage points of winter performance. The total improvement is about 40% in two years.

But the model is still far from human-level performance in snow. At 30% of its summer performance in the worst conditions, it's not safe. At 50% of its summer performance in moderate snow, it's not reliable.

The industry needs to accelerate its winter testing. It needs to collect more winter data, develop better winter perception algorithms, and train the models to understand the physics of snow.

How to test: a practical guide for OEMs

  1. Drive the worst days—after a storm, in freezing rain, at night, when the salt spray is thickest.

  2. Collect data in multiple locations—Michigan, Minnesota, Colorado, New York, not just California.

  3. Dedicate a test team to winter conditions, not just a few winter test days.

  4. Test on the same routes in summer and winter to measure the degradation.

  5. Measure perception confidence in winter conditions. Not just collision avoidance.

  6. Test in simulation and on real roads, and compare the performance. The simulation-to-real gap is huge in winter.

The Timing Log

0 entries · timing stand

No observations filed for this run yet.

Log an observation

Course observers & crew — file what you saw at the trap. Entries are stamped as witnessed.