Score
Designs, builds, and evaluates interactive conversational systems and agents that manage multi-turn text or spoken dialogue, including intelligent customer-service assistants. Work covers components such as intent and entity understanding, dialog/state management, response generation and selection, context tracking and grounding, evaluation and safety, and integration with backend services and delivery channels.
Large language model–driven conversational systems face significant engineering challenges in scaling across social and commercial applications. Method: This paper pioneers the establishment of “Conversational Systems Engineering” (CSE) as a formal knowledge domain, grounded in the SWEBOK 4.0 framework, and systematically maps its end-to-end technical lifecycle via systematic literature review and domain analysis—identifying critical research gaps in requirements modeling, evaluation & validation, and continuous operations. Contribution/Results: First, it formally positions CSE as an emerging subfield of software engineering. Second, it introduces the first structured, extensible knowledge classification framework for conversational systems. Third, it delineates dedicated engineering methodologies tailored to conversational AI. Collectively, these contributions provide a theoretical foundation for academic research and a methodological roadmap for industrial practice.
This study addresses the weak interactive capability of large language models (LLMs) in multi-turn dialogue. Methodologically, it introduces a unified “capability–evaluation–enhancement–evolution” analytical framework—the first to jointly characterize context retention, dynamic response generation, and interactive modeling. The approach integrates dialogue state tracking, long-context modeling, trajectory-driven evaluation metrics, retrieval-augmented generation, and memory mechanisms, yielding a scalable multi-turn evaluation protocol and a collaborative interactive evolution pathway. The work provides the first systematic survey of LLM interactivity, precisely identifying core bottlenecks—including state drift, long-range forgetting, and evaluation misalignment—and offers both theoretical foundations and practical guidelines for applications such as conversational search, intelligent consulting, and interactive pedagogy.
Domain-specific chatbots suffer from ambiguous user intent, contextual fragmentation, and interaction disorganization during multi-turn interactions—such as conditional filtering, multi-option selection, and comparative operations—due to the absence of GUI-like “submit/reset” mechanisms. To address this, this work introduces, for the first time, a form-based Submit/Reset paradigm into conversational systems, explicitly modeling user confirmation behaviors and context-switching actions. Methodologically, we integrate formalized state tracking, fine-grained user action recognition, and chain-of-thought (CoT) reasoning, augmented by prompt engineering to enhance large language models’ capacity for structured dialogue state representation. Experiments in hotel booking and customer management domains demonstrate significant improvements: +28.6% in multi-turn task coherence, +32.1% in user satisfaction, and a reduction of 2.4 turns on average, indicating enhanced operational efficiency.
This study investigates whether explicit intent recognition is a necessary prerequisite for generating high-quality responses in task-oriented dialogue systems. Method: Challenging the conventional intent-module-dependent paradigm, we propose and comparatively evaluate two strategies—“intent-first” and “end-to-end direct generation”—using large language models (e.g., T5) fine-tuned on multiple public task-oriented dialogue datasets. Evaluation encompasses both linguistic quality and task completion rate. Contribution/Results: Experiments demonstrate that, in typical service scenarios, direct generation achieves performance on par with or exceeding that of the intent-first approach—even without intent annotations—while substantially reducing system complexity and inference latency. These findings empirically challenge the assumed necessity of explicit intent recognition, providing evidence and conceptual support for lightweight, low-latency service assistant design grounded in a new end-to-end paradigm.
This paper addresses the lack of a systematic evaluation framework for large language model (LLM)-driven multi-turn dialogue agents. Following the PRISMA methodology, we systematically review 250 studies to propose a dual-dimensional taxonomy—“what to evaluate” and “how to evaluate.” Methodologically, we introduce a novel five-dimensional evaluation framework encompassing task completion, response quality, user experience, memory retention, planning capability, and tool utilization. We further categorize evaluation approaches into four types: human annotation, automated metrics (e.g., BLEU, ROUGE), human-AI collaboration, and LLM-based self-evaluation—the first such classification in the literature. The resulting structured knowledge system constitutes the first comprehensive, principled evaluation framework specifically designed for multi-turn dialogue agents. It establishes a unified benchmark for empirical assessment and supports scalable, multi-paradigm research in conversational AI.
Current large language models (LLMs) rely on utterance-level understanding in speech interaction, resulting in high response latency and disjointed dialogue flow—failing to meet the stringent real-time requirements of human-robot interaction (HRI). To address this, we propose a method integrating incremental automatic speech recognition (ASR), streaming language generation, partially observable Markov decision processes (POMDPs), and reactive planning to realize end-to-end word-level response generation. We formally define the core requirements and evaluation dimensions of incremental dialogue management (DM) for the first time, and establish architectural principles and practical constraints tailored to embodied robotic platforms. Our analysis identifies critical gaps in the incremental capabilities of existing DM frameworks and constructs a cross-module incremental interaction technology taxonomy. This work delivers the first deployable design guideline and implementation framework for real-time HRI systems, enabling responsive, fluid, and contextually grounded spoken dialogue.
This work addresses the challenge that existing testing methodologies struggle to effectively validate the behavior of conversational AI systems under interactive and multi-granularity integration scenarios. To this end, it proposes the first hierarchical testing framework specifically designed for conversational AI, systematically covering verification across multiple levels—from language and AI components, through single-agent systems, to multi-agent configurations. The approach integrates principles from software testing theory, architectural analysis of dialogue systems, and AI behavior validation techniques to construct a layered testing strategy. Experimental results demonstrate that the framework significantly enhances system reliability and consistency across different integration layers, offering a scalable and structured testing paradigm for complex conversational AI systems.
This study addresses the challenge of generating personalized interaction strategies for sales-oriented dialogue agents. Motivated by the finding that occupational attributes exert the strongest influence on user dialogue intent—compared to age or gender—we propose a lightweight, occupation-aware interaction strategy guidance mechanism. Our approach introduces a persona-informed user simulator framework that jointly integrates persona-sensitive policy modeling with intent-priority scheduling, circumventing computationally expensive multi-attribute coupling. The method achieves significant efficiency gains without compromising effectiveness: experiments demonstrate a 21.3% reduction in average dialogue turns and a 16.7% improvement in conversion rate. By decoupling occupational signals from other demographic features, our framework enables interpretable, low-overhead personalization—establishing a novel paradigm for scalable, intent-driven sales dialogue systems.
Existing large language models inadequately simulate authentic multi-turn human user behaviors—such as informal phrasing, personalized stylistic expression, and real-time self-correction—leading to biased and unrealistic evaluations of assistant models. Method: We propose User Language Models (User LMs), trained via human-centric post-training on multi-turn dialogue data—not by naïvely inverting assistant models—to explicitly capture realistic user interaction patterns. Contribution/Results: User LMs significantly improve behavioral fidelity and evaluation robustness, validated through both automated metrics and human assessments. Experiments reveal that when evaluated using User LMs, GPT-4o’s accuracy drops from 74.6% to 57.4% on programming and mathematical reasoning tasks, uncovering previously masked interaction weaknesses. This work establishes the first systematic, trustworthy user-side simulation paradigm for dialogue evaluation, introducing a new benchmark for assessing conversational capabilities of large models.
High-quality dialogue data is critical for task-oriented dialogue systems, yet human annotation is prohibitively expensive; it remains unclear whether LLM-generated synthetic dialogues can reliably substitute for real human interactions. Method: We introduce the first multidimensional user behavior analysis framework for task-oriented dialogue—spanning strategy, interaction style, and evaluation—and conduct parallel human–agent dialogue collection and quantitative comparison across four representative scenarios, employing behavioral dimension modeling and controlled experiments. Contribution/Results: Our analysis reveals systematic agent-user biases in feedback polarity, linguistic style, and hallucination awareness; however, agents closely match human users in problem-solving effectiveness and search strategy. This work establishes a reproducible, scalable analytical paradigm and empirical benchmark for LLM-based user simulation in task-oriented dialogue research.