🤖 AI Summary
This study addresses the inefficiency in retrieved information delivery caused by existing fixed search interfaces, which constrain agents' control over candidate processing and evidence presentation. We propose Programmatic Search Agents that treat locally executable computations as atomic search actions, transcending traditional query-rewriting limitations. By constructing persistent workspaces, generating program units, and resolving runtime data dependencies, our approach enables flexible primitive composition and selective evidence presentation. Furthermore, it supports incremental policy adaptation during training-free inference across multiple backbone models. Evaluations on the InfoSeek-Eval and BrowseComp-Plus benchmarks demonstrate absolute task success rate improvements of 4.00% and 7.56%, respectively, alongside average reductions of 28.3% to 33.9% in final-step token consumption.
📝 Abstract
Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the returned evidence can reduce subsequent search. We introduce Programmatic Search Agent (PSA), which makes a local executable computation over candidates the unit of a search action. PSA unifies a persistent candidate workspace, flexible primitive composition, and selective evidence presentation. It incrementally generates program cells that reuse candidates, execute dependent operations, and select what the agent inspects next. The runtime resolves specified data dependencies within each cell, while the agent adapts its search strategy across cells as new evidence arrives. We compare PSA with the Query-based Agent and Tool-based Agent on InfoSeek-Eval and BrowseComp-Plus using five policy backbones without task-specific training. All three interfaces share the search substrate, and the Tool-based Agent also shares PSA's primitives and persistent workspace. Relative to the Query-based Agent, PSA improves macro-averaged task success by 4.00 and 7.56 percentage points on the two benchmarks, respectively; within-backbone reductions in final-step tokens average 28.3% and 33.9%. These results support extending agent control beyond query reformulation to the processing and presentation of retrieved evidence. Code will be released subject to approval.