🤖 AI Summary
This work investigates the impact of input distribution shift and parameter quantization on the spectral properties of the empirical Fisher Information Matrix (FIM), with a focus on the behavior of its largest eigenvalue. Leveraging Weyl’s inequality, assumptions from differential geometry, matrix perturbation theory, and statistical modeling, the authors derive two-sided theoretical bounds for the largest eigenvalue under structured perturbations and propose a computationally tractable approximation method with formal guarantees. Extensive experiments across 12 models and 1,080 training trajectories confirm that quantization substantially inflates the largest eigenvalue—reaching up to 244 times the full-precision baseline under 4-bit quantization—aligning closely with theoretical predictions and revealing the pronounced disruptive effect of low-bit quantization on the FIM’s spectral structure.
📝 Abstract
We study the spectral perturbation of the empirical Fisher Information Matrix (FIM) of a parametric statistical model under two structured perturbations: departure of the input from a reference (in-distribution) ensemble, and finite-precision (quantized) perturbation of the model's parameters. For the first, under an explicit local curvature-monotonicity hypothesis on the dominant eigenvalue lambda_max of the FIM, we show departure from a reference manifold provably elevates lambda_max relative to a calibration baseline (Proposition 3.2), and discuss why this hypothesis is required, since curvature need not increase monotonically under every perturbation. Our principal result is a directional eigenvalue perturbation bound, via Weyl's inequality, showing lambda_max under a quantization noise perturbation is lower bounded by its unperturbed value up to a third-order remainder, and, under a mild genericity condition, strictly exceeds it at leading order (Theorem 4.3). We give two tractable approximations to lambda_max -- one heuristic, one with a rigorous two-sided bound -- and a completeness result for a threshold-based partition of an augmented state space. These results motivate using sigma_t = lambda_max(F_t)/lambda_base as a runtime monitoring statistic for deployed language models: the quantization result offers a mechanism for an empirical observation of our own, where a calibration threshold for this statistic was approximately 244 times larger than a preliminary full-precision estimate on a 4-bit quantized model, a single measurement rather than a value derived in closed form. We report supporting measurements (twelve models, n=1,080 trajectories) broadly consistent with our predictions, discuss the scope and limitations of every result, and state as an open problem the closed-form prediction of the quantization inflation magnitude our bound does not supply.