Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistent performance gains in post-training quantization of large language models caused by poorly understood weight averaging mechanisms. We formulate weight averaging as a trade-off between preserving training progress and enhancing robustness. By deriving a unified averaging kernel and constructing a continuous family of averaging kernels to approximate the Pareto frontier, we establish a theoretically grounded framework applicable across diverse quantization precisions. Experiments validate both the accuracy of our theoretical predictions and the effectiveness of the proposed strategies. This work provides a solid theoretical foundation and practical methodology for optimizing quantization transitions. The associated code has been made publicly available.
📝 Abstract
Large language models (LLMs) are typically pretrained in high precision but increasingly deployed with low-precision post-training quantization (PTQ). Recent studies have shown that using weight averaging during pretraining can improve PTQ performance compared with learning-rate decay, suggesting that it might provide a simple way to improve the pretraining-to-quantization transition. But the mechanism behind weight averaging remains insufficiently explained. This leads to inconsistent and fragile performance gains, thereby preventing practitioners from applying such a technique confidently. As a response, we formulate weight averaging as a trade-off between retaining training progress and improving robustness under perturbation. We further derive a continuous family of averaging kernels that unifies conventional strategies and achieves the Pareto frontier between the two competing goals. Critically, a theoretical framework for performing weight averaging under PTQ is developed. It can be shown that coarser quantization is more susceptible to perturbations, whereas finer quantization could be less affected. Thus, our results could provide unified theoretical guidance for performing weight averaging under different PTQ conditions. Experiments validate both the predicted behavior and the proposed averaging strategy. Code is available at https://github.com/MOFA-LAB/weight-averaging-for-ptq.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Post-Training Quantization
Weight Averaging
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weight Averaging
Post-Training Quantization
Large Language Models
Averaging Kernels
Perturbation Robustness
💼 Related Jobs
No related jobs found.
H
Hanzhang Wang
City University of Hong Kong
T
Tianqi Shen
City University of Hong Kong
Z
Zonglin Liu
City University of Hong Kong
J
Junze He
City University of Hong Kong
Difan Zou
Difan Zou
The University of Hong Kong
Machine LearningDeep LearningOptimizationStochastic AlgorithmsSignal Processing
Ziye Ma
Ziye Ma
Assistant Professor, CS, City University of Hong Kong
OptimizationMachine LearningEstimation