🤖 AI Summary
This study addresses the lack of theoretical grounding and questionable empirical evidence regarding the relationship between loss landscape flatness and generalization. Leveraging a teacher-student tree committee machine model, we analytically characterize the empirical risk minimization estimator and its Hessian spectrum in the high-dimensional limit. By integrating the zero-temperature Gibbs formalism, the Edwards-Jones derivation, and finite-size simulations, we construct an analytically tractable theoretical framework. Our analysis reveals that the flatness-generalization relationship depends critically on task type and the parameter-to-data ratio. In regression tasks, the spectral mean and right edge correlate stably with generalization error. Conversely, in classification tasks, positive correlation emerges only in the strongly overparameterized regime, while underparameterization even induces a reversal. These findings refine prevailing assumptions that posit a universally monotonic link between flatness and generalization.
📝 Abstract
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical minimizers of the empirical loss. Secondly, we use Edwards-Jones formalism to derive the limiting Hessian resolvent around these typical minimizers. All predictions agree with finite-size gradient-descent simulations. Finally, we study three measures of flatness, namely the left and right edges and the spectral mean, and check if a decrease in generalization error as the dataset size is increased corresponds to an increase in flatness. We find that the answer strongly depends on the learning task and on the ratio of the number of parameters to the number of data points. In regression, the spectral mean and right edge correlate with the generalization error, while the left edge does so only in the overparametrized regime. In classification this correlation reliably holds only in the highly overparametrized phase, while for underparametrized networks it can even reverse.