Information-Geometric Decomposition of Generalization Error in Unsupervised Learning

📅 2026-04-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the decomposition of generalization error in unsupervised learning by leveraging information geometry. For the first time, it rigorously decomposes the KL generalization error into three non-negative components—model error, data bias, and variance—using the generalized Pythagorean theorem and a dual e-mixture variance identity. Under an ε-PCA model with isotropic Gaussian assumptions, closed-form expressions for each component are derived, revealing that the optimal feature retention threshold coincides precisely with the noise level ε. The analysis further constructs a phase diagram delineating three distinct regimes—retain-all, interior, and collapse—that characterize phase transitions in the error structure. Comprehensive numerical experiments validate the theoretical predictions.

Technology Category

Machine Learning: Unsupervised & Self-Supervised LearningReasoning under Uncertainty: Graphical ModelsNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
We decompose the Kullback--Leibler generalization error (GE) -- the expected KL divergence from the data distribution to the trained model -- of unsupervised learning into three non-negative components: model error, data bias, and variance. The decomposition is exact for any e-flat model class and follows from two identities of information geometry: the generalized Pythagorean theorem and a dual e-mixture variance identity. As an analytically tractable demonstration, we apply the framework to $ε$-PCA, a regularized principal component analysis in which the empirical covariance is truncated at rank $N_K$ and discarded directions are pinned at a fixed noise floor $ε$. Although rank-constrained $ε$-PCA is not itself e-flat, it admits a technical reformulation with the same total GE on isotropic Gaussian data, under which each component of the decomposition takes closed form. The optimal rank emerges as the cutoff $λ_{\mathrm{cut}}^{*} = ε$ -- the model retains exactly those empirical eigenvalues exceeding the noise floor -- with the cutoff reflecting a marginal-rate balance between model-error gain and data-bias cost. A boundary comparison further yields a three-regime phase diagram -- retain-all, interior, and collapse -- separated by the lower Marchenko--Pastur edge and an analytically computable collapse threshold $ε_{*}(α)$, where $α$ is the dimension-to-sample-size ratio. All claims are verified numerically.
Problem

Research questions and friction points this paper is trying to address.

generalization error
unsupervised learning
information geometry
Kullback-Leibler divergence
model decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

information geometry
generalization error decomposition
unsupervised learning
ε-PCA
phase diagram
G
Gilhan Kim
Department of Physics and Astronomy, Seoul National University, Seoul 08826, Republic of Korea