Towards Unbiased On-Policy Distillation for Block Diffusion Language Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training instability and overconfidence collapse encountered in block diffusion models during large-block-size distillation by proposing the Un-OPD framework. This framework introduces a novel boundary-aware step filtering mechanism alongside a support rebalancing calibration strategy, which effectively eliminate optimization bias while preventing context misalignment and premature overconfidence collapse. Furthermore, by integrating online policy distillation with a block diffusion language model architecture and rollout reuse techniques, Un-OPD substantially reduces computational overhead. Experimental results demonstrate that the proposed method yields significant performance improvements on mathematical reasoning and code generation benchmarks while reducing training time by approximately 50%.
📝 Abstract
On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexplored. In this work, we investigate this regime and reveal two critical optimization biases that induce severe training instability. First, mismatched block boundaries between teacher and student cause \textbf{\textit{context misalignment}}, providing distorted supervisory signals that misguide student decoding. Second, even under aligned contexts, an \textbf{\textit{intrinsic optimization bias}} in OPD, where the student tends to rapidly absorb high-support signals while lagging on low-support updates, drives a premature confidence surge that traps weaker students in catastrophic overconfidence collapse. To resolve these, we propose \mbox{\textbf{Un-OPD}}, an unbiased on-policy distillation framework with two novelties for stabilizing BDLM training. First, Un-OPD introduces a boundary-aware step filtering strategy that eliminates context-misaligned decoding steps. Second, Un-OPD proposes moderating optimization intensity at high-support positions via a support-rebalanced confidence calibration, thereby bypassing overconfidence collapse. Beyond stability, we also introduce a rollout reuse mechanism to reduce rollout generation overhead. Extensive experiments on math reasoning and code generation benchmarks show that Un-OPD consistently stabilizes training and delivers superior performance while reducing wall-clock training time by approximately half.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Block Diffusion Language Models
Context Misalignment
Optimization Bias
Overconfidence Collapse
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Block Diffusion Language Models
Context Misalignment
Confidence Calibration
Rollout Reuse
🔎 Similar Papers
No similar papers found.
Z
Zaiquan Yang
City University of Hong Kong
Fei Wei
Fei Wei
Research Scientist, Alibaba Group
network information theoryprivacysecuritylearning
Y
Yong Wang
Alibaba Group
Y
Yudong Han
Beijing Institute of Technology
Y
Yiyu Li
City University of Hong Kong
Zhuofan Zong
Zhuofan Zong
MMLab, The Chinese University of Hong Kong
Large ModelsMultimodalObject Detection3D Object Detection
G
Gerhard Petrus Hancke
City University of Hong Kong
X
Xiangxiang Chu
Alibaba Group
R
Rynson WH Lau
City University of Hong Kong (Dongguan)