Whitening Improves Robustness to Spurious Correlations in Linear Probes

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of linear probes to spurious correlations, which arises from their inherent bias toward directions associated with large eigenvalues and consequently degrades generalization performance. To mitigate this issue, the paper establishes a theoretical connection between linear probing and maximum-margin classifiers, proposing a covariance whitening strategy that equalizes the eigenvalue spectrum and thereby eliminates directional biases. Notably, this approach operates as a general-purpose preprocessing technique that requires neither prior knowledge nor labeled data. Experiments on both synthetic datasets and standard benchmarks demonstrate that the proposed whitening strategy significantly enhances the robustness of linear probes against spurious correlations while effectively improving the performance of existing methods.
📝 Abstract
Deep neural networks tend to rely on simple features that may be spurious and thus fail to generalize. We study this problem in the setting of linear probes, where a (generalized) linear model is fitted on the representations of a (pretrained) model. We use the connection of these models to the max-margin classifier, and show they favor directions associated with large eigenvalues of the covariance matrix. Whitening removes this preference by equalizing the eigenvalues of the covariance matrix. This observation motivates whitening as a preprocessing step that can reduce reliance on spurious correlations without requiring prior knowledge of their presence or labeled data. We examine the effect of whitening on a synthetic data-generating process and standard spurious correlation benchmarks, and find that it improves robustness. We also find that whitening can improve robustness when added to existing approaches.
Problem

Research questions and friction points this paper is trying to address.

spurious correlations
linear probes
robustness
deep neural networks
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Whitening
Spurious Correlations
Linear Probes
Robustness
Max-margin Classifier
Floris Holstege
Floris Holstege
University of Amsterdam
Interpretable Machine LearningOut-of-distribution generalisation
B
Bram Wouters
University of Amsterdam, Department of Quantitative Economics; Tinbergen Institute
N
Noud van Giersbergen
University of Amsterdam, Department of Quantitative Economics; Tinbergen Institute
C
Cees Diks
University of Amsterdam, Department of Quantitative Economics; Tinbergen Institute