🤖 AI Summary
This paper systematically characterizes overfitting behavior along the full regularization path of two-part-code Minimum Description Length (MDL)-based learning rules for binary classification. Adopting the agnostic PAC framework and asymptotic analysis, it derives, for the first time, an explicit limiting expression for the generalization error as a function of the regularization parameter λ and noise level, and establishes tight worst-case upper bounds. The main contributions are threefold: (1) correcting and significantly strengthening GL’s earlier conclusion on non-uniformity at λ = 1; (2) revealing that under-regularization risk exhibits a continuous spectrum—rather than a binary threshold—across λ; and (3) proving that overfitting severity can be continuously and controllably tuned by λ, refuting the conventional “phase-transition” view. Collectively, these results provide the first comprehensive theoretical characterization and quantitative control principle for MDL regularization across the entire λ domain.
📝 Abstract
We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, based on an arbitrary prior or description language. citet{GL} previously established the lack of asymptotic consistency, from an agnostic PAC (frequentist worst case) perspective, of the MDL rule with a penalty parameter of $lambda=1$, suggesting that it underegularizes. Driven by interest in understanding how benign or catastrophic under-regularization and overfitting might be, we obtain a precise quantitative description of the worst case limiting error as a function of the regularization parameter $lambda$ and noise level (or approximation error), significantly tightening the analysis of citeauthor{GL} for $lambda=1$ and extending it to all other choices of $lambda$.