Score
Designs, builds, or evaluates conversational systems and components that manage extended interactions across multiple turns between participants. This includes modeling and tracking dialogue state and context/history, turn-taking and grounding, history-aware response generation, intent and slot tracking, clarification and error-recovery strategies, and dialogue policies that preserve coherence and continuity over an entire session.
This study addresses the weak interactive capability of large language models (LLMs) in multi-turn dialogue. Methodologically, it introduces a unified “capability–evaluation–enhancement–evolution” analytical framework—the first to jointly characterize context retention, dynamic response generation, and interactive modeling. The approach integrates dialogue state tracking, long-context modeling, trajectory-driven evaluation metrics, retrieval-augmented generation, and memory mechanisms, yielding a scalable multi-turn evaluation protocol and a collaborative interactive evolution pathway. The work provides the first systematic survey of LLM interactivity, precisely identifying core bottlenecks—including state drift, long-range forgetting, and evaluation misalignment—and offers both theoretical foundations and practical guidelines for applications such as conversational search, intelligent consulting, and interactive pedagogy.
This work addresses the evaluation and enhancement of large language models (LLMs) in multi-turn interactive settings across realistic domains—including mathematics, programming, healthcare, education, and adversarial jailbreaking—where key challenges involve long-horizon contextual consistency, response robustness, and fairness. We propose the first multidimensional benchmark taxonomy specifically designed for multi-turn dialogue, unifying three technical paradigms: intrinsic model capabilities, external augmentations (e.g., retrieval, memory, knowledge graphs), and agent-level coordination. We construct a structured evaluation resource repository and open-source an extensible challenge framework alongside practical guidelines (Awesome-Multi-Turn-LLMs). Our contributions provide a standardized benchmark, a systematic methodology, and reproducible baselines—advancing rigorous, comparable, and scalable research on multi-turn LLM interaction.
This study addresses core challenges in task-oriented dialogue with large language models (LLMs), including weak topical coherence, insufficient knowledge progression, inconsistent role embodiment, and coarse-grained controllability. To this end, we propose the first systematic, multi-dimensional parametric framework for dialogue quality control. The framework defines nine quantifiable and intervenable control parameters across six dimensions—semantic coherence, knowledge evolution, role consistency, among others—enabling fine-grained, reproducible modeling and regulation of dialogue attributes. Empirical evaluation on mainstream LLMs demonstrates statistically significant improvements in dialogue quality (p < 0.01) and task adaptability. The framework supports diverse application scenarios, including education, psychological counseling, customer service, and entertainment. By establishing a standardized, parameter-driven paradigm for dialogue generation quality control, this work advances controllable, reliable, and domain-adaptable conversational AI.
Domain-specific chatbots suffer from ambiguous user intent, contextual fragmentation, and interaction disorganization during multi-turn interactions—such as conditional filtering, multi-option selection, and comparative operations—due to the absence of GUI-like “submit/reset” mechanisms. To address this, this work introduces, for the first time, a form-based Submit/Reset paradigm into conversational systems, explicitly modeling user confirmation behaviors and context-switching actions. Methodologically, we integrate formalized state tracking, fine-grained user action recognition, and chain-of-thought (CoT) reasoning, augmented by prompt engineering to enhance large language models’ capacity for structured dialogue state representation. Experiments in hotel booking and customer management domains demonstrate significant improvements: +28.6% in multi-turn task coherence, +32.1% in user satisfaction, and a reduction of 2.4 turns on average, indicating enhanced operational efficiency.
Contemporary dialogue systems typically adopt an integrated architecture combining large language models (LLMs), external tools, and databases; thus, evaluating only the underlying LLM fails to ensure end-to-end quality. Existing evaluation methods predominantly focus on single-turn analysis and lack automated, process-aware testing for full conversational trajectories. Method: We propose the first end-to-end testing framework based on non-cooperative user simulation: (1) a challenging, role-driven user simulator requiring no reference dialogues or system-internal knowledge; (2) a fine-grained error taxonomy to guide prompt optimization and enhance detection of dialogue failures and anomalies; and (3) a decoupled architecture enabling low-cost configuration and cross-system portability. Contribution/Results: Experiments demonstrate substantial improvements in defect detection rates, with strong generalizability, scalability, and robustness across diverse dialogue systems and evaluation settings.
This study systematically investigates the persistent performance degradation of role-playing large language models (LLMs) in ultra-long dialogues (>100 turns), focusing on dynamic decay across three dimensions: role fidelity, instruction adherence, and safety. We introduce the first dialogue-conditioned long-horizon evaluation protocol, benchmarking seven prominent open- and closed-source models using long-context modeling and multi-dimensional dynamic quantitative metrics. Our analysis reveals, for the first time, a fundamental long-term trade-off between role fidelity and instruction adherence: all models exhibit significant erosion of role consistency as dialogue length increases—particularly in goal-directed scenarios—where responses progressively converge toward role-agnostic baselines, confirming a structural failure in long-term role persistence. These findings expose an intrinsic fragility in current role-playing paradigms and establish a reproducible benchmark with actionable insights for developing trustworthy, long-interaction role-aware LLMs.
Existing personalized dialogue systems struggle to model the dynamic evolution of users’ latent states, often relying on static profiles or explicit memory and lacking proactive decision-making capabilities for future interactions. This work proposes PUMA, a novel framework that introduces the free energy principle to dialogue personalization by formulating the task as a partially observable sequential decision-making process. PUMA represents user states through latent variables and integrates action-conditioned state transitions with Bayesian belief updating, guiding dialogue policy by minimizing expected free energy. This approach shifts the paradigm from passive response retrieval to active, state-evolution-driven decision-making, unifying cognitive exploration with task-oriented objectives. Experimental results demonstrate that PUMA significantly improves long-term dialogue performance on healthcare consultation and motivational interviewing datasets, achieving superior response quality, user state estimation, and next-state prediction.
This study addresses the insufficient reliability of large language models (LLMs) in real-world multi-turn, cross-topic dialogues and the absence of a systematic evaluation framework. We propose the first benchmark specifically designed to assess LLM robustness under practical interactive challenges, comprising three core tasks: maintaining cross-topic constraints, selecting appropriate tools under mixed intents, and tracking structured entities amid distractors. Through controlled single-turn versus multi-turn experiments, we conduct stress tests on leading open-source and commercial models. Our analysis uncovers critical failure modes—including instruction drift, intent confusion, and context overshadowing—and demonstrates significant performance degradation across all models in multi-turn settings, particularly among smaller-scale architectures. These findings provide essential empirical grounding and actionable insights for the trustworthy deployment and future improvement of conversational AI systems.
This work addresses the limited explicit control over dialogue acts in large language models (LLMs), which hinders their ability to reliably generate specific conversational behaviors such as confirmation or questioning. To overcome this, the authors propose the Latent-IM framework, which preserves the end-to-end architecture of LLMs while decoupling dialogue act selection from utterance generation, thereby recovering capabilities akin to traditional systems in state estimation and action control. By integrating contextual modeling, causal generation control, and latent-space intervention, Latent-IM enables plug-and-play guidance of dialogue acts. Experimental results demonstrate that, on a human dialogue act reproduction task, the method improves end-to-end dialogue act accuracy by 12.5 percentage points over an unguided baseline, achieving performance comparable to fine-tuned approaches.
Current evaluation methodologies for dialogue systems are constrained by static scoring criteria and fixed scenarios, limiting their ability to capture dynamic behaviors in multi-turn interactions. This work proposes an adaptive co-evolutionary evaluation framework that integrates generation and assessment into an iterative optimization loop through a closed-loop collaboration between a dialogue planner and a reflective analyzer. The framework employs structured templates to guide a user simulator in generating goal-oriented dialogues and automatically refines evaluation rubrics based on behavioral pattern analysis. This approach simultaneously enhances test case complexity and diagnostic precision of scoring, significantly improving coverage and accuracy in evaluating the multi-turn capabilities of advanced dialogue systems while reducing reliance on manual intervention.