🤖 AI Summary
This study addresses the limitation that existing activation quantization methods for diffusion models fail to exploit the strong correlation between the conditional and unconditional branches in classifier-free guidance (CFG), thereby constraining generation fidelity. To overcome this, we propose Guided Correlated Branch Transformation (GCBT), which derives an orthogonal matrix offline to rotationally align both branches, enabling their joint compression. This method analytically solves for the optimal basis of each layer in closed form, eliminating the need for gradient-based optimization or angle search. As a plug-and-play module, GCBT can be seamlessly integrated into existing post-training quantization frameworks. It significantly enhances image generation fidelity while preserving the underlying architecture and incurring no inference performance degradation.
📝 Abstract
Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating CFG structure into diffusion quantization, activation quantization still operates independently across conditional and unconditional coordinates, leaving cross-activation structure unexploited. We show that matched CFG activations form a strongly correlated two-dimensional source and that, under a fixed bit budget, the choice of branch coding basis materially affects quantization fidelity. Motivated by this observation, we introduce branch-space transform coding, which rotates matched CFG branches via an offline derived 2x2 orthogonal matrix, requiring minimal modifications to model parameters or the quantization pipeline. We further derive the Guidance-Correlation Branch Transform (GCBT), which jointly incorporates the CFG guidance direction and cross-branch second moments. Under an equal-rate quantization-noise surrogate, GCBT admits a closed-form per-layer solution without gradient optimization or angle search. Applied on top of existing diffusion PTQ methods, GCBT yields statistically significant fidelity gains in most evaluated comparisons with no statistically significant degradation, while leaving the underlying host quantization pipeline unchanged.