Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the generalization bias in post-training quantization of large vision-language models (LVLMs) caused by an over-reliance on reconstruction loss. To overcome this limitation, we propose Balanced Fitting, a framework that departs from the conventional error-minimization paradigm by exploiting the regularization benefits that quantization confers upon specific layers and modalities. Through fine-grained evaluation of component-wise quantization effects, a hybrid fitting strategy, and joint weight-activation quantization, our approach dynamically balances accuracy preservation with regularization gains. Extensive experiments demonstrate that the proposed method significantly outperforms existing baselines across diverse LVLM architectures. These findings compellingly establish that low reconstruction loss does not necessarily translate to superior downstream performance, thereby introducing a new paradigm for multimodal model quantization.
📝 Abstract
Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss with respect to the full-precision model, potentially over-preserving FP behavior and calibration-specific bias. Rather than treating quantization solely as an error to be minimized, we observe that it can also provide beneficial regularization for certain layers and modalities. Motivated by this observation, we propose Balanced Fitting, a quantization effect-based framework that balances precision and regularization beyond reconstruction-based optimization. By measuring layer- and component-wise quantization effects for weights, vision activations, and text activations, Balanced Fitting combines fine-grained fitting for sensitive components with coarser fitting to exploit potential regularization benefits. Experiments on multiple LVLMs show that our method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance. The source code is publicly available at https://github.com/kmc3661/BFQ
Problem

Research questions and friction points this paper is trying to address.

Post-Training Quantization
Large Vision-Language Models
Reconstruction Loss
Generalization
Regularization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-training quantization
Large vision-language models
Balanced fitting
Regularization
Reconstruction loss
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Minchan Kang
Korea Advanced Institute of Science and Technology (KAIST)
K
Kyeonghye Park
Korea Advanced Institute of Science and Technology (KAIST)
S
Seungyeon Sa
Korea Advanced Institute of Science and Technology (KAIST)
S
Seoyoung Cho
Korea Advanced Institute of Science and Technology (KAIST)
D
Daeshik Kim
Korea Advanced Institute of Science and Technology (KAIST)
Y
Yucheol Cho
Hanbat National University