Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge for pivotal commitments in masked diffusion language model training by proposing an offline self-distillation framework. The method introduces information-gain-based selection of "pivot" tokens to apply differentiated, precise supervision across successful and failed trajectories. By integrating cross-entropy loss, directed unlikelihood estimation, and uncertainty quantification, the framework optimizes only high-impact commitment points, thereby circumventing the inefficiency of full-sequence training. Experimental results demonstrate that this approach substantially enhances the reasoning performance of LLaDA-8B with minimal data, outperforming both supervised fine-tuning and compute-matched reinforcement learning baselines.
📝 Abstract
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Masked diffusion language models
Credit assignment
Self-distillation
Denoising
Post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Masked Diffusion Language Models
Credit Assignment
Information Gain
Unlikelihood Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Seo Hyun Kim
Seo Hyun Kim
KAIST
S
Sunwoo Hong
KAIST AI, University of Toronto & Vector Institute
Y
Younwoo Choi
University of Toronto & Vector Institute
C
Chen-Hao Chao
University of Toronto & Vector Institute
S
Se-Young Yun
KAIST AI
Rahul G. Krishnan
Rahul G. Krishnan
University of Toronto
Machine LearningArtificial IntelligenceHealthcareProbabilistic modelsCausal Inference