SoloQ: Calibration-Free Quantization for Diffusion Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that quantizing diffusion language models typically requires calibration data, as activation distributions vary drastically during denoising. To overcome this, we propose a calibration-free quantization framework that introduces a structured K-RPBH rotation combined with a lightweight rescaling mechanism. This approach maps both weights and activations onto a predictable normalized rotational basis, enabling data-independent low-bit inference via distribution-matching codebooks. Furthermore, the framework supports hardware-native NVFP4 formats and KV cache compression for block diffusion models. Experimental results demonstrate that under 4-bit quantization, our method surpasses calibration-dependent baselines in accuracy while reducing peak memory consumption by 2.61× and accelerating end-to-end inference speed by 2.24×.
📝 Abstract
Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation quantization is therefore attractive, yet existing dLLM post-training quantization methods rely on calibration data despite activation distributions shifting across masking states and denoising steps. We present SoloQ, a calibration-free quantization framework that maps weights and activations into a normalized rotated basis with a predictable marginal distribution, enabling data-independent quantization. SoloQ combines a structured K-RPBH rotation with a lightweight rescaling correction for calibration-free quantization. Its predictable post-rotation distribution supports both distribution-matched codebooks and hardware-native NVFP4. For block-diffusion models, SoloQ further applies commit-time KV-cache quantization to compress persistent states without perturbing the actively denoised block. Across full-sequence dLLMs (LLaDA and Dream) and block-diffusion dLLMs(Fast-dLLM v2 and Nemotron-Labs-Diffusion), SoloQ retains accuracy under 4-bit quantization and outperforms calibration-based baselines on knowledge- and reasoning-intensive benchmarks. With NVFP4, SoloQ reduces peak memory by up to 2.61X and accelerates end-to-end inference by up to 2.24X.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Language Models
Quantization
Calibration-Free
Inference Efficiency
KV-cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

Calibration-Free Quantization
Diffusion Language Models
K-RPBH Rotation
Commit-time KV-cache Quantization
NVFP4
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Donghyun Lee
University of Southern California
Arkapravo Ghosh
Arkapravo Ghosh
Yale University
Deep LearningAI Accelerator
V
Varun Manjunath
University of Southern California
B
Bumjoon Kyle Rhee
University of Southern California
H
Hyunho Kook
University of Southern California
S
Shiting Xiao
Yale University
Youngeun Kim
Youngeun Kim
Applied Scientist, Amazon AWS AI Labs
Machine LearningEfficient AINeuromorphic Computing
P
Priyadarshini Panda
University of Southern California