What Does the Driver Represent?
The problem
The first challenge is deciding what counts as an interpretable internal signal. Activation Atlases established a useful intuition: meaningful visual concepts may live in patterns of jointly activated units rather than in individual neurons, and those patterns can expose spurious correlations that ordinary accuracy metrics conceal. (openai.com)
Anthropic’s feature-based work sharpens the same idea in a different setting. Its argument is not that individual neurons cleanly encode concepts, but that sparse, distributed features can provide a more useful unit of analysis. That distinction matters for driving models, where “yielding,” “closing gap,” or “emergency vehicle nearby” may be distributed across time, sensors, and interacting agents rather than localized in one activation.
Wayve’s scenario-intelligence proposal brings the question into autonomous driving. Its clustered representations appear to organize scenes around behaviorally relevant factors, such as following distance, while abstracting over lighting, weather, and vehicle appearance. The promise is practical: use the model’s own representation space to inspect data coverage and find rare cases. The unresolved issue is whether such clusters are stable semantic concepts, task-specific shortcuts, or merely convenient partitions of the dataset. (wayve.ai)
Readings