Trajectory-Retrieval Speculative Decoding: When Does a Model's Own History Help?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive decoding costs associated with long chain-of-thought reasoning. We propose Trajectory Locally Adaptive Retrieval (TLAR), a method that repurposes a model's historical trajectories as runtime memory to accelerate generation. TLAR introduces a novel adaptive retrieval mechanism that dynamically adjusts the activation threshold and candidate tree width, coupled with an exact verification algorithm. This combination optimizes speculative decoding while strictly preserving output distribution consistency. Experimental evaluations demonstrate that TLAR substantially improves token acceptance rates and end-to-end throughput across code debugging, mathematical reasoning, and writing tasks, achieving a significant breakthrough in inference efficiency.
📝 Abstract
Long chain-of-thought reasoning increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history supplies useful drafts and complements an existing drafter. Controlled source comparisons reveal trajectory-specific reuse, motivating our method Trajectory-Local Adaptive Retrieval (TLAR). TLAR retrieves approximately matched continuations from the current trajectory and uses recent verification outcomes to adapt retrieval activation and candidate width. TLAR combines retrieved continuations with model-generated drafts in a shared candidate tree, preserving the target model's output distribution through exact verification. Across code debugging, mathematics, and open-ended writing, our evaluation connects source reuse, incremental acceptance, and execution cost. Combining TLAR with strong retrieval baselines improves token acceptance under matched verification budgets and increases end-to-end throughput over the draft-model baseline. These findings support generated trajectories as runtime memory for adaptive inference.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
chain-of-thought reasoning
trajectory retrieval
decoding efficiency
draft reuse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Trajectory-Local Adaptive Retrieval
Chain-of-Thought Reasoning
Candidate Tree
Adaptive Inference
🔎 Similar Papers
No similar papers found.