TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human Preference

📅 2025-06-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address challenges in multi-turn preference elicitation with large language models (LLMs)—including redundant questioning, inefficient dialogue trajectories, and generation of task-irrelevant queries—this paper proposes the first dialogue trajectory optimization framework. It decouples clarification path planning from response summarization to enable task-driven dynamic questioning and semantically coherent summarization. Methodologically, it integrates reinforcement learning–based trajectory optimization, a task-conditioned clarification parser, and a controllable summarization module, jointly modeling user preferences and multi-step reasoning. On standard preference elicitation benchmarks, our approach achieves a 9.32% absolute improvement over strong baselines, reduces irrelevant query generation by 37.6%, and significantly enhances user feedback adoption and task completion accuracy. The core contribution is a trajectory–response co-optimization paradigm that explicitly formulates dialogue path planning as a learnable, controllable, task-aligned process—the first such formulation in the literature.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Learning Preferences or RankingsKnowledge Representation and Reasoning: Preferences

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) can effectively elicit human preferences through multi-turn dialogue. Complex tasks can be accomplished through iterative clarifying questions and final responses generated by an LLM acting as a questioner (STaR-GATE; Andukuri et al., 2024}). However, existing approaches based on self-taught reasoning struggle to identify optimal dialogue trajectories and avoid irrelevant questions to the tasks. To address this limitation, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization, which consists of two key components: a clarification resolver that generates optimal questioning trajectories, and a summarizer that ensures task-aligned final responses. The trajectory optimization enables the model to produce effective elicitation questions and summary responses tailored to specific tasks. Experimental results demonstrate that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation tasks.
Problem

Research questions and friction points this paper is trying to address.

Optimizing dialogue trajectories for human preference elicitation
Reducing irrelevant questions in multi-turn LLM dialogues
Improving task-aligned question generation and response summarization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory optimization for dialogue enhancement
Clarification resolver for optimal questioning
Summarizer for task-aligned responses
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yulin Dou
School of Information Science and Engineering, Yunnan University, China; Yunnan Key Laboratory of Intelligent Systems and Computing, China
Jiangming Liu
Jiangming Liu
Associate Professor, Yunnan University
natural language processingdeep learning