🤖 AI Summary
This study addresses the vulnerability of linear probes to spurious correlations, which arises from their inherent bias toward directions associated with large eigenvalues and consequently degrades generalization performance. To mitigate this issue, the paper establishes a theoretical connection between linear probing and maximum-margin classifiers, proposing a covariance whitening strategy that equalizes the eigenvalue spectrum and thereby eliminates directional biases. Notably, this approach operates as a general-purpose preprocessing technique that requires neither prior knowledge nor labeled data. Experiments on both synthetic datasets and standard benchmarks demonstrate that the proposed whitening strategy significantly enhances the robustness of linear probes against spurious correlations while effectively improving the performance of existing methods.
📝 Abstract
Deep neural networks tend to rely on simple features that may be spurious and thus fail to generalize. We study this problem in the setting of linear probes, where a (generalized) linear model is fitted on the representations of a (pretrained) model. We use the connection of these models to the max-margin classifier, and show they favor directions associated with large eigenvalues of the covariance matrix. Whitening removes this preference by equalizing the eigenvalues of the covariance matrix. This observation motivates whitening as a preprocessing step that can reduce reliance on spurious correlations without requiring prior knowledge of their presence or labeled data. We examine the effect of whitening on a synthetic data-generating process and standard spurious correlation benchmarks, and find that it improves robustness. We also find that whitening can improve robustness when added to existing approaches.