Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of mean-accuracy-based configuration selection in serving block diffusion models, which cannot guarantee acceleration without compromising reference accuracy. To overcome this, we propose Redline, a method that replaces conventional mean-based baselines with a distribution-free, finite-sample risk control mechanism. By evaluating operating-point risks on a calibration set, Redline deploys the fastest configuration satisfying a user-specified failure budget. The approach integrates self-distillation, static skipping rules, speculative decoding acceptance strategies, and weight quantization to achieve quantifiably safe acceleration. Experimental results on LLaDA2 math tasks demonstrate that Redline delivers significant speedups while strictly adhering to a 10% risk budget, outperforming mean-based selection rules.
📝 Abstract
Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on prompts that the slower one answers correctly. On the serving engine and its decode traces, the default commit rule already commits every fully resolved block, a static skip rule captures nearly all of the compute that allocation can save, and self-distillation on engine-decoded targets adds speed at unchanged accuracy. Larger speedups come from lower thresholds, which commit tokens that are still uncertain. We therefore present Redline, a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts. Redline keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget with high probability, and deploys the fastest configuration that passes. It speeds up math at a smaller risk budget than code in both model families, and at a budget of ten percent it deploys a LLaDA2 math configuration that commits over a third more tokens in each forward. It also applies without modification to the acceptance rule of speculative decoding and to weight quantization. On the same calibration data, Redline stays within its stated failure probability, whereas each tolerance of a mean-accuracy rule either gains less speed for some model and task or exceeds the risk budget far more often for another. Code is available at https://github.com/js-lee-AI/Redline.
Problem

Research questions and friction points this paper is trying to address.

block-diffusion language models
serving acceleration
risk guarantees
operating point selection
speculative decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Block-diffusion models
Distribution-free risk guarantees
Speculative decoding
Finite-sample procedure
LLM serving optimization
🔎 Similar Papers
No similar papers found.