Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of information leakage awareness, capability confounding, and benchmark scarcity in evaluating disambiguation for tabular conversational systems. To this end, we propose a fuzzy verifiable task framework that introduces a novel formal definition of leakage awareness to decouple inquiry from resolution strategies. This framework unifies representations across six datasets and establishes the AmbiTab benchmark. By integrating reinforcement learning with a judge-free leakage diagnostic metric, the approach enables precise leakage quantification and environment modeling. Experimental results demonstrate that the proposed method significantly improves disambiguation metrics across all evaluated datasets and enhances task success rates on five of them, while effectively revealing the impact of oracle leakage during training.
📝 Abstract
Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conversational Task Disambiguation
Oracle Leakage
Ambiguous Verifiable Task
Asking Policy
Reinforcement Learning
N
Nafiseh Ghoroghchian
Layer 6 AI
Luis Scoccola
Luis Scoccola
CRM-ISM - LaCIM@Université du Québec à Montréal - Université de Sherbrooke
computational topology and geometrygeometric machine learningpersistence theorytype theory
T
Tina Sedaghat
Layer 6 AI
O
Omid Vaheb
Layer 6 AI
H
Hannah Chen
Layer 6 AI
D
Dino D'Agostino
Layer 6 AI
K
Keyvan Golestan
Layer 6 AI