TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the computational redundancy in static quantization of diffusion models, which uniformly applies high-precision arithmetic across all denoising steps. The authors propose a time-adaptive bit-sparse quantization method that dynamically adjusts quantization precision per timestep and network layer for the first time. By sharing high-precision weights, learning spatiotemporal least-significant-bit masks, and integrating a temporal precision engine, the approach enables bit-serial execution without switching overhead. It decouples memory and computation costs, eliminating the need for multi-stage weight copies or runtime search. Evaluated on PixArt-Sigma, SANA-1.6B, and SDXL-Turbo, the method preserves generation quality while reducing inference latency by 25%–50% and achieving 6.1–7.5× speedup over naive 8-bit serial execution.
📝 Abstract
Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-case arithmetic throughout the denoising trajectory. We introduce Temporal-Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TASQ stores one shared maximum-precision weight buffer and learns a Temporal-Spatial LSB Mask that selects a lower effective precision for each layer and denoising stage by truncating least-significant bits. Storage therefore remains fixed by the worst case, while BitOPs decrease at less sensitive stages without per-stage weight copies or runtime search. A Temporal-Precision Engine maps the learned schedule to bit-serial execution, where cycles scale with effective precision and switching precision has no measured cycle overhead. On PixArt-Sigma, SANA-1.6B, and SDXL-Turbo, TASQ achieves quality comparable to static quantization with less computation. Together with the Temporal-Precision Engine, it reduces execution cycles by 25 to 50 percent over static quantization and by 6.1 to 7.5x over a naive static 8-bit bit-serial execution. Code is available at https://github.com/seokho-han/tasq.
Problem

Research questions and friction points this paper is trying to address.

diffusion models
quantization
temporal adaptation
bit precision
denoising steps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal-Adaptive Quantization
Bit Sparsification
Diffusion Models
Bit-Serial Execution
Precision Scheduling
S
Seokho Han
Department of Electrical and Computer Engineering, Sungkyunkwan University, Korea
D
Dongwei Wang
Department of Electrical and Computer Engineering, University of Arizona, USA
J
Jinhee Kim
Department of Electrical and Computer Engineering, Duke University, USA
Yiran Chen
Yiran Chen
John Cocke Distinguished Professor of Electrical and Computer Engineering, Duke University
memoryneuromorphicmachine learning systems
K
Kang Eun Jeon
Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)
Huanrui Yang
Huanrui Yang
Assistant Professor, ECE, University of Arizona
Efficient deep learningTrustworthy deep learning
Jong Hwan Ko
Jong Hwan Ko
SungKyunKwan Univ. (SKKU)
Deep learning acceleratorImage/audio processingVLSI/IoT systems design