Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of conventional post-training quantization methods for multimodal large language models, which neglect autoregressive feedback and consequently induce future state drift. To mitigate this, we propose OnPTQ, a framework that performs online calibration over quantization policy trajectories. We introduce a novel decision-consequence risk mechanism that combines instantaneous discrepancies with short-horizon counterfactual rollouts to evaluate critical decision boundaries, prioritizing states that significantly impact generation. Furthermore, context anchoring and trajectory refreshing techniques are incorporated to achieve precise calibration. Extensive experiments on the Qwen model series demonstrate that OnPTQ substantially improves downstream performance under low-bit settings and effectively reduces correctness flips relative to FP16 baselines, all without altering the inference graph.
πŸ“ Abstract
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.
Problem

Research questions and friction points this paper is trying to address.

Post-training quantization
Multimodal large language models
Autoregressive feedback
On-policy calibration
Decision-consequence risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-training quantization
On-policy calibration
Multimodal large language models
Counterfactual rollouts
Decision-Consequence risk
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
W
Wenxiao Fan
School of Computer Science and Technology, Beijing Institute of Technology
J
Jingling Fu
JD.com
L
Lichen Ma
JD.com; Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
Y
Yu He
JD.com
L
Luohang Liu
JD.com
J
Jinbao Xue
JD.com
K
Ke Zhang
JD.com
Junshi Huang
Junshi Huang
Meituan
Computer VisionNLPMachine Learning
Kan Li
Kan Li
Huazhong University of Science and Technology
3D AssemblyStretchable ElectronicsMetamaterials