When Flat Minima Fail: Characterizing INT4 Quantization Collapse After FP32 Convergence

📅 2026-04-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study challenges the common assumption that convergence in FP32 implies suitability for quantization by revealing a sharp degradation in INT4 performance after FP32 convergence. Analyzing all 154 training checkpoints of Pythia-160M using calibration-free per-group INT4 probing, the authors identify a three-phase explosive growth in quantization error post-convergence and, for the first time, link the onset of this collapse to FP32 perplexity convergence rather than learning rate decay. They propose Oscillatory Lock-In, a novel learning rate schedule combined with kurtosis-based outlier filtering, which significantly enhances INT4 robustness. In multi-schedule comparisons, this approach effectively mitigates the surge in late-stage quantization error—from 11% to 517%—reducing it on average by 2.2 percentage points (p<0.0001), thereby validating the critical role of schedule amplitude calibration.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Post-training quantization (PTQ) assumes that a well-converged model is a quantization-ready model. We show this assumption fails in a structured, measurable, and previously uncharacterized way. Using a calibration-free per-group INT4 probe applied to all 154 publicly available Pythia-160m training checkpoints, we identify a three-phase divergence structure: a rapid-learning phase where both FP32 perplexity and quantization robustness improve together, a meta-stable plateau lasting roughly 70,000 steps where FP32 perplexity stagnates but INT4 gap remains bounded, and an explosive divergence phase where the INT4 gap compounds from 11% to 517% while FP32 perplexity barely moves. Critically, this divergence begins not when the learning rate starts decaying, but precisely when FP32 perplexity converges a finer-grained onset predictor that implies post-convergence weight updates, rather than decay magnitude alone, are the proximate cause. We further show that INT8 quantization is entirely immune throughout all three phases, constraining the mechanism to the coarseness of the 16-level INT4 grid specifically, and rule out weight outlier accumulation as the mechanism via direct kurtosis measurement. Finally, we conduct a controlled fork experiment from the pre-divergence checkpoint comparing three learning rate schedules (cosine continuation, SGDR warm restarts, and our proposed Oscillatory Lock-In) across nine independent runs. SGDR uniformly accelerates divergence (0/9 pairwise wins against cosine), while OLI's settled cool phases reduce the INT4 gap by 2.2 percentage points on average (t = -5.46, p < 0.0001), demonstrating that schedule amplitude calibration, not oscillation alone, determines whether perturbation helps or hurts. Our code, probe implementation, and all 154-checkpoint audit results are released publicly.
Problem

Research questions and friction points this paper is trying to address.

INT4 quantization
quantization collapse
post-training quantization
FP32 convergence
quantization robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

INT4 quantization collapse
post-convergence divergence
calibration-free probing
learning rate schedule
quantization robustness
💼 Related Jobs
No related jobs found.
M
Marcus Armstrong
Department of Computer Science, University of Houston