🤖 AI Summary
To address challenges in multi-turn preference elicitation with large language models (LLMs)—including redundant questioning, inefficient dialogue trajectories, and generation of task-irrelevant queries—this paper proposes the first dialogue trajectory optimization framework. It decouples clarification path planning from response summarization to enable task-driven dynamic questioning and semantically coherent summarization. Methodologically, it integrates reinforcement learning–based trajectory optimization, a task-conditioned clarification parser, and a controllable summarization module, jointly modeling user preferences and multi-step reasoning. On standard preference elicitation benchmarks, our approach achieves a 9.32% absolute improvement over strong baselines, reduces irrelevant query generation by 37.6%, and significantly enhances user feedback adoption and task completion accuracy. The core contribution is a trajectory–response co-optimization paradigm that explicitly formulates dialogue path planning as a learnable, controllable, task-aligned process—the first such formulation in the literature.
📝 Abstract
Large language models (LLMs) can effectively elicit human preferences through multi-turn dialogue. Complex tasks can be accomplished through iterative clarifying questions and final responses generated by an LLM acting as a questioner (STaR-GATE; Andukuri et al., 2024}). However, existing approaches based on self-taught reasoning struggle to identify optimal dialogue trajectories and avoid irrelevant questions to the tasks. To address this limitation, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization, which consists of two key components: a clarification resolver that generates optimal questioning trajectories, and a summarizer that ensures task-aligned final responses. The trajectory optimization enables the model to produce effective elicitation questions and summary responses tailored to specific tasks. Experimental results demonstrate that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation tasks.