Mentored Decoding: Faster Inference meets Boosting

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between inference speed and generation quality in lossy speculative decoding. By integrating speculative decoding with Boosting theory, this work proposes a "Mentor Decoding" framework that unifies optimization objectives through f-divergence generalization and introduces a divergence-agnostic O(n) space data structure for efficient sampling. Notably, it provides the first formal proof that lossy speculative decoding can surpass the quality upper bound of the target model, revealing deep connections between inference processes and classical training theory. Algorithmically, the proposed method achieves O(n) distribution construction and O(log n) optimal parameter querying, yielding simultaneous breakthroughs in both inference acceleration and generation quality.
📝 Abstract
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored decoding}$. To get there, we connect inference to a celebrated ML training theory, $\textit{boosting}$, and proceed via the generalization of mentored decoding to the whole set of $f$-divergences. We uncover key properties of mentored decoding, among which (i) the particularly appealing geometric nature of the total variation case, (ii) simple approximations for any $f$-divergence in direct relation with boosting compliance, and (iii) a $\textit{divergence independent}$ $O(n)$ space and $O(\mathrm{sort}(n))$ time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem in $O(\log n)$ time and constructing optimal mentored distributions in $O(n)$ time for any $f$-divergence.
Problem

Research questions and friction points this paper is trying to address.

Speculative Decoding
Mentored Decoding
Boosting
f-divergences
Inference Acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mentored Decoding
Speculative Decoding
Boosting
f-divergences
Inference Acceleration
💼 Related Jobs
No related jobs found.
V
Vivien Tran-Thien
Google
Richard Nock
Richard Nock
Google Research