BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between the training objectives of diffusion draft models and the sequence-level block verification mechanism by proposing a block-verification-aware loss function. This method is the first to directly derive the inference-stage block verification rules into a training objective, achieving end-to-end sequence-level optimization by maximizing the expected acceptance length, thereby effectively aligning the training and inference processes. Experiments on Qwen3 models demonstrate that the proposed approach increases the average number of accepted tokens by 13%–21%, significantly outperforming cross-entropy and existing token-level loss baselines. These results establish a superior training paradigm for diffusion-based speculative decoding.
📝 Abstract
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatch, we introduce Block Verification-aware loss (BV loss), a training objective designed to maximize the expected acceptance length of a drafted sequence. BV loss is directly derived from the block verification acceptance rule, providing a principled connection between the drafter training objective and the inference-time verification mechanism at the sequence level. Across math, code, and chat benchmarks, BV loss increases the mean number of tokens accepted per verification call under block verification by 13.0--21.0\% over cross-entropy loss training for DFlash and DSpark with Qwen3-4B and Qwen3-8B without changing the inference procedure. BV loss also outperforms tokenwise acceptance objectives such as TV loss and LK loss, and its gains extend to token verification and greedy decoding. These results demonstrate the benefit of training block diffusion drafters with an objective aligned with sequence-level verification, rather than optimizing each token independently.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
diffusion drafter
block verification
training objective mismatch
sequence-level drafting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Block Verification-aware Loss
Diffusion Drafter
Sequence-level Verification
Training Objective Alignment
S
Suyoung Kim
a2sys (a2sys.ai); Seoul National University (snu.ac.kr)
J
Jahyun Koo
a2sys (a2sys.ai); Seoul National University (snu.ac.kr)
H
Hyeonjin Kim
a2sys (a2sys.ai); Seoul National University (snu.ac.kr)
I
Inhyeok Bang
a2sys (a2sys.ai)
S
Seunghyun Lee
a2sys (a2sys.ai)
H
Hyunjae Oh
a2sys (a2sys.ai); University of Wisconsin (wisc.edu)
B
Baeseong Park
a2sys (a2sys.ai)
Dongsoo Lee
Dongsoo Lee
NAVER Cloud
Model compressionoptimizationAI Chip Design