dialogue state tracking

Designs, builds, or analyzes components that estimate user intent and slot values and maintain a structured representation of the dialogue state across turns. This includes updating that state from incoming (including multimodal) inputs, grounding and tracking entity references, resolving cross-turn anaphora, and exposing the maintained state for downstream response generation.

dialoguestatetracking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges

May 19, 2025
HW
Hongru Wang
🏛️ The Chinese University of Hong Kong | The University of Edinburgh | Macquire University | Georg-August Universität Göttingen | Université de Montréal | MILA | Beihang Univeristy

Existing tool-use evaluation benchmarks predominantly focus on single-turn, stateless scenarios, neglecting the dynamic state evolution and tool lifecycle dependencies inherent in multi-turn dialogues. This work introduces DialogTool—the first benchmark explicitly designed for stateful, multi-turn tool usage—alongside VirtualMobile, a configurable virtual mobile environment. DialogTool systematically models the full tool lifecycle across three phases: tool creation, perception/selection/execution, and role-consistent response generation. Its innovations include multi-turn state-tracking annotations, an API simulation execution environment, and a six-task, phase-wise evaluation protocol. We conduct a comprehensive cross-model evaluation across 13 state-of-the-art LLMs. Results reveal significant performance degradation in long-horizon, state-sensitive tool use: all models exhibit consistent accuracy decay across dialogue turns, exposing fundamental limitations in state representation and long-range dependency modeling.

Assessing stateful tool use in multi-turn dialoguesEvaluating whole tool lifecycle across six key tasksTesting LLM robustness in long-horizon tool interactions

This study addresses the issue in multi-turn dialogues where rejected or replaced historical intents continue to interfere with large language model decision-making, leading to task failure. To tackle the model's difficulty in distinguishing between "mentioned" and "active" intents, this work proposes the Intent-Eval benchmark and the Intent-OPSD framework. Built upon online policy self-distillation, the framework initializes teacher and student networks from a homogeneous model and employs decision-conditioned supervision with a frozen teacher to precisely guide the model toward focusing on currently active intents. Experimental results demonstrate that the proposed method significantly enhances robustness in scenarios such as tool calling and code generation, effectively mitigating performance degradation caused by irrelevant historical dialogue turns.

intent trackinglarge language modelsmentioned-as-in-effect confusion

Existing personalized dialogue systems struggle to model the dynamic evolution of users’ latent states, often relying on static profiles or explicit memory and lacking proactive decision-making capabilities for future interactions. This work proposes PUMA, a novel framework that introduces the free energy principle to dialogue personalization by formulating the task as a partially observable sequential decision-making process. PUMA represents user states through latent variables and integrates action-conditioned state transitions with Bayesian belief updating, guiding dialogue policy by minimizing expected free energy. This approach shifts the paradigm from passive response retrieval to active, state-evolution-driven decision-making, unifying cognitive exploration with task-oriented objectives. Experimental results demonstrate that PUMA significantly improves long-term dialogue performance on healthcare consultation and motivational interviewing datasets, achieving superior response quality, user state estimation, and next-state prediction.

dialogue personalizationlatent user statesmulti-turn conversation

Feedstack: Layering Structured Representations over Unstructured Feedback to Scaffold Human AI Conversation

Jun 03, 2025
HN
H. Nguyen
🏛️ Temple University | University of San Diego | Berea College

Current human–AI dialogue feedback systems lack mechanisms to support deep reflection and shared understanding between users and AI. Method: This study proposes a hierarchical interaction architecture that overlays structured representations onto unstructured dialogue streams, enabling users to organize, navigate, and externalize feedback. It innovatively integrates the design probe paradigm to enhance exploratory and reflective capabilities, and combines research-driven design, hierarchical interface modeling, and feedback structural encoding in two formative user studies (n = 16). Contribution/Results: The approach significantly improves novice designers’ accuracy in articulating design intent and their ability to identify core design principles. Findings provide both a novel conceptual pathway and empirical grounding for developing next-generation dialogue feedback systems that are explainable, traceable, and supportive of collaborative cognition.

Enhancing feedback conversations with structured layered representationsInvestigating interface design for organizing and externalizing feedbackScaffolding exploration and shared understanding in human-AI dialogue

Existing user simulators struggle to precisely control turn-level intents, often leading to deviations from the intended dialogue goals. This work proposes decoupling user intent from language generation by introducing, for the first time, an explicit turn-level intent interface that enables fine-grained control over user behavior through instruction-conditioned generation. Built upon the UserIDA framework, the approach integrates supervised fine-tuning with population-based reinforcement learning to calibrate intents effectively, substantially enhancing the simulator’s goal-directed capabilities. Experimental results demonstrate that the proposed method achieves an intent accuracy of 86.6%, representing a 24.3-percentage-point improvement over the strongest baseline, and successfully fulfills at least four distinct goal intents in 91.7% of dialogue states.

controllable generationdialogue systemsintent control

Latest Papers

What's happening recently
View more

This study addresses the failure mode in which LLM-based agents, during multi-turn interactions, suffer degraded output quality due to historical instructions interfering with evolving user intents. We formally define and quantify this phenomenon as "intent drift," and construct IntentFlux, an executable benchmark designed to evaluate it. To mitigate this issue, we propose StateForge, a method that explicitly tracks and maintains active requirement states to suppress interference from obsolete instructions. Experimental results demonstrate that intent drift significantly degrades task performance, while StateForge effectively alleviates this degradation by improving the average score from 0.367 to 0.467. This work provides a principled state-maintenance mechanism for robust multi-turn dialogue systems.

Intent DriftLLM AgentsMulti-turn Interaction

This study addresses the lack of explicit state representations in large language models for task-oriented dialogue, which undermines reliable information updating and behavior. By investigating how internal dialogue states are encoded, we identify a dissociation between structural and value-based linear readabilities. Leveraging this finding, we propose a novel structure-readout-based state-action controller paradigm that selectively rectifies model actions via structural readouts without requiring full belief state prediction. Validated through linear probing, causal interventions, and closed-loop evaluations across multiple datasets, this approach improves exact query accuracy from 0.318 to 0.621 and task success rate from 0.272 to 0.371 across five instruction-tuned models, while incurring minimal computational overhead.

behavioral reliabilitybelief stateconversational state

Large language models excel in static, single-turn tasks but struggle to effectively track and adapt to the evolving nature of user intent in dynamic, multi-turn dialogues. To address this limitation, this work proposes an evaluation framework that reformulates static tasks as dynamic multi-turn conversations, incorporating an intent evolution simulation mechanism that gradually reveals, revises, or shifts user intent throughout the interaction while remaining compatible with existing evaluation protocols. This framework systematically uncovers a performance gap in current models under dynamic intent scenarios, demonstrating significant discrepancies between static and dynamic settings across multiple tasks. The findings highlight a critical shortcoming in models’ ability to support collaborative interaction and establish a reusable paradigm for dynamic evaluation in future research.

dynamic interactionevolving user intentintent tracking

This study addresses the challenge of generating personalized interaction strategies for sales-oriented dialogue agents. Motivated by the finding that occupational attributes exert the strongest influence on user dialogue intent—compared to age or gender—we propose a lightweight, occupation-aware interaction strategy guidance mechanism. Our approach introduces a persona-informed user simulator framework that jointly integrates persona-sensitive policy modeling with intent-priority scheduling, circumventing computationally expensive multi-attribute coupling. The method achieves significant efficiency gains without compromising effectiveness: experiments demonstrate a 21.3% reduction in average dialogue turns and a 16.7% improvement in conversion rate. By decoupling occupational signals from other demographic features, our framework enables interpretable, low-overhead personalization—establishing a novel paradigm for scalable, intent-driven sales dialogue systems.

Adapting dialogue strategies based on user demographic profilesAutomating personalized interaction planning for conversational agentsEnhancing sales-oriented dialogue effectiveness through persona-informed strategies

Hot Scholars

GT

Gokhan Tur

University of Illinois Urbana-Champaign
Conversational AILanguage UnderstandingLarge Language Models
DH

Dilek Hakkani-Tür

Professor of Computer Science, Univ. Illinois Urbana-Champaign
Speech and Language ProcessingDialogue SystemsSpoken Language UnderstandingMachine Learning
RX

Ruifeng Xu

Professor, Harbin Institute of Technology at Shenzhen
Natural Language ProcessingAffective ComputingArgumentation MiningLLMs
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning
WH

Wenyi Hong

Tsinghua University
multimodal pretraining