arXiv · 2609.39177
Whitening Improves Robustness to Spurious Correlations in Linear Probes
Abstract
Deep neural networks tend to rely on simple features that may be spurious and thus fail to generalize. We study this problem in the setting of linear probes, where a (generalized) linear model is fitted on the representations of a (pretrained) model. We use the connection of these models to the max-margin classifier, and show they favor directions associated with large eigenvalues of the covariance matrix. Whitening removes this preference by equalizing the eigenvalues of the covariance matrix. This observation motivates whitening as a preprocessing step that can reduce reliance on spurious correlations without requiring prior knowledge of their presence or labeled data. We examine the effect of whitening on a synthetic data-generating process and standard spurious correlation benchmarks, and find that it improves robustness. We also find that whitening can improve robustness when added to existing approaches.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Floris Holstege, Bram Wouters, Noud van Giersbergen, Cees Diks. 2026-09-30. Whitening Improves Robustness to Spurious Correlations in Linear Probes. https://arxiv.org/abs/2609.39177
Cite the original work for its findings. Save a collection to share your selection of sources.