Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the substantial accuracy degradation and inefficient low-rank compensation encountered when applying 4-bit quantization to multimodal diffusion Transformers. We propose a unified post-training quantization framework that formulates low-rank-assisted W4A4 quantization as a coupled calibration problem. Methodologically, we introduce the concept of a contracted Hessian to discount residual errors and incorporate an activation noise proxy to suppress quantization artifacts, enabling highly efficient optimization at extremely low ranks through joint calibration. Experimental results demonstrate that the proposed method surpasses SVDQuant across multiple architectures using only rank-4 approximations. Furthermore, it accelerates the quantization process by up to 6.25× while significantly reducing computational resource consumption, establishing a highly efficient paradigm for compressing large-scale generative models.
📝 Abstract
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating the output-side low-rank factor yields a \emph{deflated Hessian} that discounts residual errors already captured by the low-rank component, while an activation-noise surrogate is incorporated to suppress activation quantization error. Across five diffusion backbones, rank-4 \method{} consistently outperforms rank-4 SVDQuant in PSNR and LPIPS. It further surpasses rank-32 SVDQuant on SANA-1.6B, FLUX.1-schnell, and FLUX.1-dev with an $8\times$ smaller rank and up to $6.25\times$ faster quantization. Furthermore, on the Qwen3-8B LLM, rank-4 \method{} improves MMLU accuracy from 61.50\% to 68.17\% over rank-32 SVDQuant. Overall, \method{} achieves better W4A4 performance with substantially lower rank and quantization cost.
Problem

Research questions and friction points this paper is trying to address.

post-training quantization
diffusion transformers
W4A4
low-rank compensation
activation quantization error
Innovation

Methods, ideas, or system contributions that make the work stand out.

W4A4 Quantization
Post-Training Quantization
Deflated Hessian
Low-Rank Compensation
Diffusion Transformers