🤖 AI Summary
This study addresses the problem that quantization errors in video diffusion Transformers are nonlinearly amplified during denoising, rendering local reconstruction insufficient for predicting final degradation. To overcome this, we propose a 4-bit post-training quantization method integrating trajectory sensitivity with activation geometry. Innovatively, it leverages short-term propagation errors to predict final latent errors, surpassing conventional local reconstruction objectives. The approach estimates propagation risk via isolated block-step interventions, performs residual correction through neighborhood code editing, and optimizes row-radius selection guided by offline calibration. Experiments demonstrate that our method significantly improves key consistency and dense reference metrics on models such as Wan, validating its effectiveness.
📝 Abstract
Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry to guide offline calibration. Isolated block--step interventions estimate propagation risk, which prioritizes sensitive trajectory states during row-radius selection. With these radii fixed, response-subspace correction uses neighboring-code edits to reduce residual components along dominant activation directions. Both stages preserve the original 4-bit weight representation. Controlled interventions show that short-horizon propagated error predicts final latent error more reliably than immediate block-output error, supporting calibration beyond local reconstruction objectives. Evaluations on Wan models, Self Forcing, and MiniMax-H3 demonstrate improvements in key consistency and dense-reference metrics while remaining competitive on other attributes across model scales and generation paradigms.