AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of sparse retrieval corpora and fixed draft lengths in retrieval-based speculative decoding for coding agents. To overcome these challenges, we propose an optimization framework integrating multi-source retrieval with dynamic length adaptation. Methodologically, we introduce a pioneering hybrid retrieval strategy combining session, workspace, and global contexts to enrich candidate corpora, alongside a feedback-driven online length adaptation mechanism that dynamically adjusts draft lengths using offline profiling. Experimental results demonstrate that our approach achieves up to a 4.76× throughput improvement over autoregressive decoding, significantly outperforming existing baselines and effectively accelerating code generation efficiency.
📝 Abstract
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37$\times$ at batch size 1 and 4.76$\times$ at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
Problem

Research questions and friction points this paper is trying to address.

speculative decoding
coding agents
retrieval-based drafting
draft length adaptation
agent pipelines
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Decoding
Retrieval-Based Drafting
Coding Agents
Dynamic Draft Length
Multi-level Corpus
🔎 Similar Papers
No similar papers found.