Rational Clarification by Assistive Agents via Value-of-Information Reasoning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how language-augmented agents should balance the safety and efficiency of direct execution versus clarification when encountering ambiguous requests. To this end, this work proposes REVOIR, a training-free framework grounded in inference-time reasoning that integrates reinforcement learning expected rewards with information-theoretic information gain via a value-of-information mechanism. By dynamically evaluating the utility of asking questions, REVOIR overcomes the limitations of conventional strategies that neglect downstream performance and fail to accommodate proactive user corrections. Experimental results on ambiguous question answering and household task planning demonstrate that the proposed approach significantly improves success rates with minimal clarification queries. Furthermore, it achieves a 13–15% increase in user preference satisfaction compared to fine-tuned baselines, enabling efficient, adaptive, and precise agent services.
📝 Abstract
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.
Problem

Research questions and friction points this paper is trying to address.

assistive agents
ambiguous requests
clarification decision
value of information
uncertainty reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Value-of-Information
Inference-time Reasoning
Clarification Policy
Assistive Agents
Preference Alignment