FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models

📅 2026-06-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical vulnerability in diffusion-based large language models (dLLMs) during post-training quantization, where fragile decisions at the write frontier—due to “stability lag”—are prone to being erroneously locked by quantization error, leading to performance degradation. The study is the first to identify this frontier decision fragility and introduces a two-stage calibration framework that avoids end-to-end diffusion inference. It first constructs a positional prior using a full-precision teacher model, integrating both frontier hit likelihood and mask reliability; then performs layer-wise off-policy calibration via reweighted hidden-state mean squared error, prioritizing protection of critical frontier states. Theoretical analysis shows this weighted objective effectively approximates the KL divergence of the output distribution. Experiments demonstrate consistent performance gains across multiple benchmarks, with significant improvements over existing methods on LLaDA and Dream (W4A4), effectively suppressing decision flipping and post-commitment mismatches.
📝 Abstract
Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, leading to a "stability lag" where early decisions remain fragile even after being written. We reveal that Post-Training Quantization (PTQ) error easily flips these borderline decisions at the write frontier, which are then permanently locked in and amplified. To address this, we propose Frontier-Aware Instability-Reweighted Calibration (FAIR-Calib), a two-stage PTQ framework for dLLMs. Stage I probes a full-precision teacher to estimate a position prior that combines frontier hits and masked-stage reliability. Stage II performs off-policy, layer-wise calibration by minimizing a reweighted hidden-state MSE, effectively prioritizing the protection of fragile frontier states without requiring expensive end-to-end diffusion rollouts. We further theoretically justify our weighted objective as a surrogate for output KL divergence. Empirically, FAIR-Calib consistently outperforms state-of-the-art baselines on LLaDA and Dream (W4A4), significantly reducing frontier decision flips and suppressing post-commit mismatches across diverse benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Large Language Models
Post-Training Quantization
stability lag
frontier decision flips
quantization error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Training Quantization
Diffusion Language Models
Frontier-Aware Calibration
Instability Reweighting
Quantization Robustness
🔎 Similar Papers
No similar papers found.