SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of multimodal capabilities and statistical instability in low-bit quantization of vision-language models (VLMs) by proposing SubRot. This method innovatively introduces channel-space feature decomposition, leveraging eigendecomposition of the empirical Fisher matrix, local Taylor expansion, and first- and second-order mixed constrained optimization to perform fine-grained rotational quantization calibration within a signed gradient subspace, ensuring cross-sample stability while preserving sign sensitivity. Experimental results demonstrate that SubRot significantly outperforms FlatQuant across five VLM benchmarks. Under W4A4 quantization, the accuracy drop is limited to within 1.4%, effectively balancing deployment efficiency with multimodal performance.
📝 Abstract
Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios and the positions of visual information, limiting statistical stability. Moreover, overly coarse aggregation through absolute values and averaging discards gradient signs and channel-wise differences, limiting the separation of modality-specific sensitivities. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural basis for capturing stable task-sensitive structures. We therefore propose SubRot, a signed gradient subspace calibration method for VLM rotation quantization. Through eigendecomposition of the empirical Fisher matrix of activation gradients, SubRot identifies a sensitive channel subspace with three properties: cross-sample stability, clear sensitivity separation, and consistent signed effects on the autoregressive loss along certain directions. Guided by a local Taylor expansion, SubRot combines signed first-order guidance along sign-stable directions with second-order constraints along the remaining sensitive directions, while retaining MSE for overall reconstruction quality. This objective steers quantization errors toward loss-decreasing directions while controlling their magnitude. Experiments on five VLMs across five benchmarks show consistent average-score improvements over FlatQuant under W4A6 and W4A4, reaching 1.4 percentage points on LLaVA-NeXT-7B. Under W4A4, average accuracy degradation from FP16 remains within 1.4 percentage points across all evaluated models, while LLaVA-v1.5-13B exceeds its FP16 average score by 0.4 percentage points.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Post-Training Quantization
Low-bit Quantization
Gradient Statistics
Modality Sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rotation Quantization
Vision-Language Models
Gradient Subspace Calibration
Post-training Quantization
Empirical Fisher Matrix
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhenhao Shang
Northwest Polytechnical University
H
Haizhao Jing
Northwest Polytechnical University
Haokui Zhang
Haokui Zhang
Northwestern Polytechnical University
Approximate nearest neighbor searchneural architecture searchdepth estimationHSI classificaion
G
Guoting Wei
Nanjing University of Science and Technology
R
Rong Xiao
Intellifusion
J
Jianqing Gao
iFLYTEK CO., LTD
Peng Wang
Peng Wang
School of Computer Science, Northwestern Polytechnical University, China
Computer VisionMachine LearningArtificial Intelligence