🤖 AI Summary
This work addresses the puzzle that deep learning models often generalize well despite perfectly fitting training data—a phenomenon known as benign overfitting—that contradicts classical statistical learning theory. Prevailing explanations attribute this behavior to model “simplicity,” yet they lack rigorous theoretical links between such simplicity and generalization performance. By integrating philosophical analysis with statistical learning theory, this study uncovers implicit assumptions underlying these explanations and argues that the vague attribution of “simplicity” to individual models creates a conceptual gap between intuitive notions and formal guarantees. The paper emphasizes the necessity of establishing provable theoretical bridges connecting the simplicity of individual interpolating models to their generalization capabilities, thereby providing a solid foundation for understanding benign overfitting.
📝 Abstract
Contemporary deep learning methods generalize well even when they fit their training data perfectly, a phenomenon known as benign interpolation. This phenomenon cannot be accounted for by classical statistical learning theory and has prompted a range of attempted new explanations in the statistics and machine learning literature. A common feature of these new proposals is an appeal to a simplicity preference among interpolating models, often presented as a form of Occam's razor. We clarify this debate for a philosophical audience and argue that this new appeal to simplicity creates an explanatory gap. The classical theory offers theorems which connect the simplicity of model classes to good generalization, thus underwriting methodological simplicity norms. The new accounts instead appeal to properties of individual models, which they interpret as a kind of simplicity. Lacking a provable connection to generalization, it is the name "simplicity" that does the work a theorem used to do, making a substantive and unargued assumption look like the application of a familiar methodological principle.