🤖 AI Summary
This study addresses the limitation of traditional retrieval-augmented generation in deriving implicit evidence from multi-source data by proposing RECAST, a framework that formulates evidence construction as a sequential decision-making process. Rather than passively retrieving information, RECAST actively deduces evidence by generating executable code through the adaptive invocation of computation, access, and synthesis tools via a routing mechanism. Methodologically, it employs RouterLM, trained with supervised fine-tuning and Group Relative Policy Optimization, for policy planning, which collaborates with CompilerLM and AnswerLM to complete the reasoning loop. Experimental results demonstrate that RECAST achieves an average success rate of 75.6% across six benchmarks, outperforming the strongest baseline by 15.9% while exhibiting substantial zero-shot generalization capabilities.
📝 Abstract
Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.