Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Self-driving cars increasingly rely on deep neural networks to achieve human-like driving 1–3 . The opacity of these black-box planners makes it challenging to accurately anticipate when they will fail 4–6 , with potentially catastrophic consequences 7–9 . Although research into interpreting these systems has surged, most of it is confined to simulations or toy setups because of the difficulty of real-world deployment 10,11 , leaving the practical utility of these techniques unknown. Here, we introduce the Concept-Wrapper Network (CW-Net), a method for faithfully explaining the behaviour of machine-learning-based planners that causally grounds their reasoning in human-interpretable concepts without sacrificing performance. We deploy CW-Net on a real self-driving car and show that the resulting explanations improve the human driver’s mental model of the vehicle, allowing them to better predict its behaviour, particularly in surprising situations. This demonstrates that explainable deep learning integrated into self-driving cars can be both understandable and useful in a realistic deployment setting. We anticipate our method could be applied to other safety-critical systems, such as autonomous drones and robotic surgeons, as well as to other architectures, such as end-to-end learning systems and vision–language–action models. Overall, our study establishes a deployment-validated pathway to interpretability for autonomous agents, which could help make them more transparent and safe.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Eoin M. Kenny, Akshay Dharmavaram, Sang Uk Lee, Tung Phan-Minh, Shreyas Rajesh, Yunqing Hu, Laura Major, Momchil S. Tomov, Julie A. Shah
- Quelle
- Nature
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0028-0836, 1476-4687
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Eoin M. Kenny, Akshay Dharmavaram, Sang Uk Lee, Tung Phan-Minh, Shreyas Rajesh, Yunqing Hu, Laura Major, Momchil S. Tomov, Julie A. Shah (2026). Explainable deep learning improves human mental models of self-driving cars. Nature. https://doi.org/10.1038/s41586-026-10950-5
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1