Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between semantic bias and high-latency execution verification encountered by large language model agents during large-scale API retrieval. To tackle this challenge, we reformulate tool retrieval as a resource-constrained sequential decision-making problem and introduce the first execution-aware agent world model, which jointly predicts success rate, latency, and utility without invoking real APIs. By integrating Monte Carlo Tree Search with an uncertainty-guided mechanism, our approach enables budget-aware planning. This work proposes a novel, cost-effective, and efficient paradigm for tool retrieval. Extensive experiments on the ToolBench I3 test set demonstrate that our method achieves an NDCG@5 of 91.40%, surpassing state-of-the-art baselines by 1.24% and establishing an optimal balance between accuracy and efficiency.
📝 Abstract
Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems. Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency. We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem. At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs. This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints. Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios. On the most challenging I3 split, it attains an NDCG@5 of 91.40\%, outperforming the state-of-the-art ToolGen (90.16\%) by 1.24\%. Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.
Problem

Research questions and friction points this paper is trying to address.

Tool Retrieval
LLM-based Agents
Budget Constraint
Accuracy-Efficiency Trade-off
API Ecosystem
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tool Retrieval
Surrogate World Model
Monte Carlo Tree Search
Budget-Aware Planning
LLM Agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zongze Wu
Beijing University of Posts and Telecommunications
Y
Yani Guo
Beijing University of Posts and Telecommunications
Runnan Li
Runnan Li
Beijing University of Posts and Telecommunications