arXiv · 2608.22155
Why Does Robustness Reduce Superposition?
Abstract
The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs. We present an empirical explanation inspired by the feature taxonomy of Ilyas et al. (2019), tracing the following chain of causalities: adversarial training abandons non-robust features, leading to fewer total features to represent, resulting in less superposition.
Explore related subjects
Keep this discovery
Adam Elimadi. 2026-08-23. Why Does Robustness Reduce Superposition?. https://arxiv.org/abs/2608.22155
Cite the original work for its findings. Save a collection to share your selection of sources.