🤖 AI Summary
This work addresses the memory bottleneck in autoregressive decoding of large language models with hybrid attention, where existing tree speculative decoding methods suffer from high verification latency due to recurrent layers and excessive transient memory overhead. The authors propose a co-design of kernel functions and runtime systems that reformulates linear attention recurrence into a closed-form tree structure, enabling parallel verification of all speculative tokens and reconstructing only the states corresponding to sampled tokens. By integrating tree-structured closed-form recurrence, lossless token-level state encoding, batch-level verification scheduling, and SGLang support, this study achieves the first efficient tree speculative decoding for hybrid attention models. Experiments demonstrate 3.4–7.7× acceleration in linear attention tree verification, 82–99× reduction in transient memory usage, up to 4.72× higher offline throughput, and online task improvements with 67.6% lower time-to-first-token (TTFT) and 49.9% reduced time-per-output-token (TPOT).
📝 Abstract
Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing tree-speculation systems are designed around the key--value caches of full-attention models. On hybrid models, they traverse recurrent layers branch by branch and materialize a full state for every proposal node, causing verification latency and transient memory to scale poorly with tree and batch sizes. We present Bole, a kernel--runtime co-design that enables efficient tree speculation for hybrid-attention LLMs. Bole transforms the linear-attention recurrence into a tree-structured closed form and realizes it with a resource-efficient GPU kernel, verifying all proposal nodes in parallel and accelerating linear-attention tree verification by 3.4--7.7$\times$. It losslessly encodes speculative state updates as token-level factors and reconstructs only the state selected after sampling, reducing transient state memory by 82--99$\times$ and freeing GPU capacity for KV caches. Its integration into SGLang, a widely deployed production LLM serving engine, couples efficient state management with a batch-wide verification budget calibrated to the complete hybrid forward. Across four models, two GPU platforms, and diverse datasets, Bole delivers up to $4.72\times$ the offline decode throughput of autoregressive decoding and up to $2.03\times$ that of the strongest tree-speculative baseline. Under online agent workloads, it reduces TTFT and TPOT by up to $67.6%$ and $49.9%$, respectively, over the strongest tree-speculative baseline.