UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of parallel draft models under high-entropy distributions caused by insufficient diversity. To this end, we propose a draft tree construction method that couples a unigram proposer with a bigram selector. The core innovation lies in a tree-native training strategy based on a renormalized KL objective, which jointly optimizes cross-entropy and KL divergence to broaden the supervision scope. This approach significantly enhances branch plausibility and overcomes bottlenecks inherent in high-entropy scenarios. Evaluated on the Qwen3 model series, our method achieves 5.84× to 6.94× speedups, comprehensively outperforming state-of-the-art baselines such as DARTree and DSpark.
📝 Abstract
Speculative decoding accelerates language model inference by verifying multiple draft tokens in a single target-model pass. Recent parallel drafters have achieved breakthrough performance in frontier production models, but their effectiveness deteriorates as the entropy of target distributions increases due to insufficient draft diversity. To overcome this bottleneck without sacrificing parallelism, we introduce UBTree, a parallel drafter that couples a Unigram proposer with a Bigram selector to construct drafting Trees. The unigram proposer is trained with the standard cross-entropy objective to generate candidate tokens independently for each position, while a lightweight bigram selector predicts transition scores between adjacent candidate pairs. Unlike the proposer, the selector is trained with a renormalized KL objective on high-temperature data. This tree-native training broadens the supervision beyond the greedy path, encouraging plausible alternative branches that improve the chance of accepting additional tokens during tree verification. Across seven standardized benchmarks with Qwen3-4B and Qwen3-8B, UBTree achieves an average speedup of $5.84$--$6.94\times$ over autoregressive decoding and outperforms DARTree in all 28 comparisons. Production-scale evaluation further demonstrates UBTree's advantage over frontier baselines such as DSpark.
Problem

Research questions and friction points this paper is trying to address.

Speculative Decoding
Parallel Drafting
Draft Diversity
High Entropy
Language Model Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Parallel Tree Drafting
Unigram-Bigram Model
Renormalized KL Objective
Tree-native Training
🔎 Similar Papers
No similar papers found.