π€ AI Summary
This study addresses the limitation of conventional post-training quantization methods for multimodal large language models, which neglect autoregressive feedback and consequently induce future state drift. To mitigate this, we propose OnPTQ, a framework that performs online calibration over quantization policy trajectories. We introduce a novel decision-consequence risk mechanism that combines instantaneous discrepancies with short-horizon counterfactual rollouts to evaluate critical decision boundaries, prioritizing states that significantly impact generation. Furthermore, context anchoring and trajectory refreshing techniques are incorporated to achieve precise calibration. Extensive experiments on the Qwen model series demonstrate that OnPTQ substantially improves downstream performance under low-bit settings and effectively reduces correctness flips relative to FP16 baselines, all without altering the inference graph.
π Abstract
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.