🤖 AI Summary
This study addresses the reliance on calibration data and the quality degradation caused by layer normalization means in existing data-free quantization methods for diffusion Transformers (DiTs). To this end, we propose CentriQ, a framework that reveals the structural causes underlying the failure of Hadamard rotations. We introduce the first calibration-free quantization mechanism combining exact mean centering with rank-1 branches, which centers activations prior to rotation and subsequently restores their means. By integrating adaptive layer normalization with a robust Lp weight fitting algorithm, CentriQ enables efficient 4-bit quantization. Experiments demonstrate that CentriQ matches the performance of SVDQuant across three DiT models while significantly outperforming existing calibration-free approaches. Notably, it achieves usable image generation quality under 2-bit activation quantization for the first time.
📝 Abstract
Diffusion transformers (DiTs) achieve state-of-the-art image generation, but their sampling cost limits deployment. Quantizing both weights and activations to 4 bits reduces this cost, yet existing methods fall short in one of two ways. Calibration-based methods are tied to a specific checkpoint and prompt distribution, whereas data-free Hadamard rotation, effective for LLMs, loses quality on DiTs. We show that this loss has a structural cause. Adaptive layer-norm conditioning adds a per-token mean to the activations, and at the widths of the evaluated DiTs, the Hadamard rotations used by data-free methods cannot spread this mean uniformly across coordinates. A single dominant direction therefore survives the rotation and sets the quantization range. We introduce CentriQ, a calibration-free quantizer that centers each token before rotation and restores the mean exactly through a rank-1 full-precision branch, so that per-token scales follow in closed form without data. Weights are fitted under a robust $\ell_p$ objective that tracks the dense mode of each group and discounts heavy tails. Across three DiTs, CentriQ matches the quality of calibrated SVDQuant at 4 bits, whereas calibration-free weight quantizers with plain per-token activation quantization collapse or degrade substantially. CentriQ outperforms the strongest calibration-free method reported to date at 2-bit weights. It is also the first calibration-free method to retain usable image quality at 2-bit activations.