Ology

Reading the Model Before the Crash

Using mechanistic interpretability and activation analysis to predict unsafe behavior in autonomous-driving models

5 sessions · 15 readings · 3 views

Curated by Ben Lorenzo

What Does the Driver Represent?session 1Correlation, Cause, and the Moment Before Failuresession 2Will the Signal Travel?session 3From Internal Evidence to a Safety Casesession 4Can Simulation Validate the Monitor?session 5

Opening

Autonomous-driving systems are increasingly being built as end-to-end models: a single learned system maps sensor streams to driving behavior, often without a clean boundary between perception, prediction, and planning. That design promises better generalization across cities and scenarios, but it also removes many of the intermediate quantities that engineers traditionally test.

Mechanistic interpretability and activation-based analysis offer a different route into these systems. Instead of asking only whether a vehicle merged, yielded, or avoided a collision, researchers can inspect the internal representations that appear before the action: clusters associated with following distance, attention to road users, uncertainty, or safety-critical scene structure. Wayve’s scenario-intelligence work, for example, treats emergent representations as a way to find coverage gaps and rare cases before they become on-road failures. (wayve.ai)

But an interpretable signal is not automatically a reliable warning signal. A representation may correlate with a failure without causing it; a saliency map may describe attention without revealing the decision mechanism; a simulator may reproduce an intervention while still misrepresenting how other road users react. Recent work on closed-loop neural simulation, controllable world models, and safety evaluation makes the problem more concrete—and more difficult. (wayve.ai)

This seminar follows one question from representation to deployment: can internal signals provide trustworthy, generalizing evidence of unsafe behavior before the vehicle acts, and what would it take to validate that claim in the real world?

Five questions worth arguing about

  1. 1Can activation structure reveal driving concepts that outcome metrics systematically miss?
  2. 2When does a failure-related representation become a warning signal rather than merely a correlated description?
  3. 3Do latent signals generalize across maneuvers, sensors, geographies, and unseen environments?
  4. 4How should internal evidence change the way autonomous-driving systems are tested and monitored?
  5. 5What evidence would justify trusting activation-based safety claims beyond simulation?

The Sessions

1Session 1start here

What Does the Driver Represent?

The problem

The first challenge is deciding what counts as an interpretable internal signal. Activation Atlases established a useful intuition: meaningful visual concepts may live in patterns of jointly activated units rather than in individual neurons, and those patterns can expose spurious correlations that ordinary accuracy metrics conceal. (openai.com)

Anthropic’s feature-based work sharpens the same idea in a different setting. Its argument is not that individual neurons cleanly encode concepts, but that sparse, distributed features can provide a more useful unit of analysis. That distinction matters for driving models, where “yielding,” “closing gap,” or “emergency vehicle nearby” may be distributed across time, sensors, and interacting agents rather than localized in one activation.

Wayve’s scenario-intelligence proposal brings the question into autonomous driving. Its clustered representations appear to organize scenes around behaviorally relevant factors, such as following distance, while abstracting over lighting, weather, and vehicle appearance. The promise is practical: use the model’s own representation space to inspect data coverage and find rare cases. The unresolved issue is whether such clusters are stable semantic concepts, task-specific shortcuts, or merely convenient partitions of the dataset. (wayve.ai)

2Session 2Correlation, Cause, and the Moment Before Failure3 readings3Session 3Will the Signal Travel?3 readings4Session 4From Internal Evidence to a Safety Case3 readings5Session 5Can Simulation Validate the Monitor?3 readings