🤖 AI Summary
This study addresses the degradation of mathematical reasoning in compressed diffusion language models caused by the mismatch between calibration states and inference trajectories during low-rank compression. To mitigate this issue, we propose a trajectory-aware low-rank objective along with the Traj-MC algorithm. Theoretically, we establish optimality conditions for low-rank approximation under the trajectory distribution. Methodologically, we design an efficient Monte Carlo sampling-based estimation algorithm, Traj-MC, which optimizes the approximation quality of partially masked states to preserve the model's reasoning capabilities. Experimental results demonstrate that, under identical compression budgets, the proposed approach significantly outperforms conventional clean calibration strategies, yielding substantial improvements across mathematical reasoning benchmarks.
📝 Abstract
Diffusion language model (dLLM) compression faces a known challenge because calibration is typically performed on clean, fully visible activations, whereas inference traverses partially masked intermediate states. For low-rank compression, this raises two questions. First, can low-rank optimality still be characterized when approximation quality is measured over trajectory-distributed states, and second, does the choice of calibration states affect mathematical reasoning preservation under compression? We address these questions by formulating a trajectory-aware low-rank objective over corruption levels and masking realizations. To estimate this objective efficiently, we propose Traj-MC, which estimates the trajectory second moment through Monte Carlo sampling and yields exact sampled-state optimality and population consistency. Under matched compression budgets, trajectory-aware calibration improves reconstruction over the generation trajectory and preserves substantially more mathematical reasoning than clean calibration on mathematical reasoning benchmarks. Our results connect trajectory-aware low-rank optimality to the reasoning capability retained after dLLM compression. Our code is available at: https://github.com/Zishan-Shao/traj-mc.git.