🤖 AI Summary
This study addresses the substantial accuracy degradation and inefficient low-rank compensation encountered when applying 4-bit quantization to multimodal diffusion Transformers. We propose a unified post-training quantization framework that formulates low-rank-assisted W4A4 quantization as a coupled calibration problem. Methodologically, we introduce the concept of a contracted Hessian to discount residual errors and incorporate an activation noise proxy to suppress quantization artifacts, enabling highly efficient optimization at extremely low ranks through joint calibration. Experimental results demonstrate that the proposed method surpasses SVDQuant across multiple architectures using only rank-4 approximations. Furthermore, it accelerates the quantization process by up to 6.25× while significantly reducing computational resource consumption, establishing a highly efficient paradigm for compressing large-scale generative models.
📝 Abstract
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating the output-side low-rank factor yields a \emph{deflated Hessian} that discounts residual errors already captured by the low-rank component, while an activation-noise surrogate is incorporated to suppress activation quantization error. Across five diffusion backbones, rank-4 \method{} consistently outperforms rank-4 SVDQuant in PSNR and LPIPS. It further surpasses rank-32 SVDQuant on SANA-1.6B, FLUX.1-schnell, and FLUX.1-dev with an $8\times$ smaller rank and up to $6.25\times$ faster quantization. Furthermore, on the Qwen3-8B LLM, rank-4 \method{} improves MMLU accuracy from 61.50\% to 68.17\% over rank-32 SVDQuant. Overall, \method{} achieves better W4A4 performance with substantially lower rank and quantization cost.