🤖 AI Summary
Tool-augmented LLM agents struggle to distinguish sources of uncertainty, hindering optimal decisions between user clarification and environmental verification. This work proposes PROUR, a framework that formulates the problem as explicit uncertainty routing among ACT, CLARIFY, and VERIFY actions, precisely differentiating goal ambiguity from missing world evidence. Methodologically, PROUR routes information sources via uncertainty decomposition and optimizes a query generator through reinforcement learning with a mode-conditioned information gain reward to acquire targeted information. Experiments demonstrate that PROUR achieves a 28.17% success rate on τ-bench, surpassing the state-of-the-art by 4.57%, while effectively reducing interaction steps and exhibiting cross-domain generalization.
📝 Abstract
Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly determining the appropriate information source for each decision. We formulate this problem as uncertainty routing among ACT, CLARIFY, and VERIFY, and propose PROUR, a proactive uncertainty routing framework. PROUR decomposes action uncertainty into disagreement across plausible user-goal interpretations, which signals user-side ambiguity, and the entropy remaining within each interpretation, which signals missing world-side evidence. To acquire information from the routed source, a query generator is trained with a mode-conditioned information-gain reward, targeting user-goal identification under CLARIFY and next-action identification under VERIFY. On $\tau$-bench, PROUR achieves 28.17% average success rate across retail and airline, outperforming the strongest prior method by 4.57% while using 2.17 fewer interaction steps. The learned policy further generalizes to stronger task agents and transactional domains of $\tau^3$-bench without retraining, demonstrating the benefit of source-aligned uncertainty resolution for proactive agents.