🤖 AI Summary
This study addresses the degradation of multimodal capabilities and statistical instability in low-bit quantization of vision-language models (VLMs) by proposing SubRot. This method innovatively introduces channel-space feature decomposition, leveraging eigendecomposition of the empirical Fisher matrix, local Taylor expansion, and first- and second-order mixed constrained optimization to perform fine-grained rotational quantization calibration within a signed gradient subspace, ensuring cross-sample stability while preserving sign sensitivity. Experimental results demonstrate that SubRot significantly outperforms FlatQuant across five VLM benchmarks. Under W4A4 quantization, the accuracy drop is limited to within 1.4%, effectively balancing deployment efficiency with multimodal performance.
📝 Abstract
Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios and the positions of visual information, limiting statistical stability. Moreover, overly coarse aggregation through absolute values and averaging discards gradient signs and channel-wise differences, limiting the separation of modality-specific sensitivities. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural basis for capturing stable task-sensitive structures. We therefore propose SubRot, a signed gradient subspace calibration method for VLM rotation quantization. Through eigendecomposition of the empirical Fisher matrix of activation gradients, SubRot identifies a sensitive channel subspace with three properties: cross-sample stability, clear sensitivity separation, and consistent signed effects on the autoregressive loss along certain directions. Guided by a local Taylor expansion, SubRot combines signed first-order guidance along sign-stable directions with second-order constraints along the remaining sensitive directions, while retaining MSE for overall reconstruction quality. This objective steers quantization errors toward loss-decreasing directions while controlling their magnitude. Experiments on five VLMs across five benchmarks show consistent average-score improvements over FlatQuant under W4A6 and W4A4, reaching 1.4 percentage points on LLaVA-NeXT-7B. Under W4A4, average accuracy degradation from FP16 remains within 1.4 percentage points across all evaluated models, while LLaVA-v1.5-13B exceeds its FP16 average score by 0.4 percentage points.