Score
Designs, builds, and evaluates interactive systems that enable back-and-forth natural language (text or speech) exchanges between users and software; this includes specifying conversational flows, dialogue management and state tracking, turn-taking, and integration with external services. Work also covers implementing and analyzing components such as natural language understanding, response generation, speech recognition/synthesis (when applicable), multimodal inputs/outputs, and task- or open-domain evaluation metrics.
Large language model–driven conversational systems face significant engineering challenges in scaling across social and commercial applications. Method: This paper pioneers the establishment of “Conversational Systems Engineering” (CSE) as a formal knowledge domain, grounded in the SWEBOK 4.0 framework, and systematically maps its end-to-end technical lifecycle via systematic literature review and domain analysis—identifying critical research gaps in requirements modeling, evaluation & validation, and continuous operations. Contribution/Results: First, it formally positions CSE as an emerging subfield of software engineering. Second, it introduces the first structured, extensible knowledge classification framework for conversational systems. Third, it delineates dedicated engineering methodologies tailored to conversational AI. Collectively, these contributions provide a theoretical foundation for academic research and a methodological roadmap for industrial practice.
Existing research on large language model (LLM)-driven multi-turn dialogue systems suffers from fragmented frameworks, heterogeneous technical approaches, and inconsistent evaluation protocols. Method: This paper introduces, for the first time, a unified taxonomy of four LLM adaptation paradigms for multi-turn dialogue—prompt engineering, instruction tuning, retrieval-augmented generation (RAG), and dialogue state modeling—while distinguishing open-domain and task-oriented dialogue along technical and bottleneck dimensions. It constructs a structured technical landscape covering architectures, benchmark datasets, and evaluation metrics, and synthesizes multi-dimensional automatic and human evaluation methodologies. Contribution/Results: The survey identifies three critical research gaps—interpretability, long-horizon consistency, and controllable interaction—and establishes a theoretically grounded, practice-oriented reference framework to guide future work in LLM-based dialogue systems.
Large language models (LLMs) exhibit unstable dialogue behavior and poor maintainability in complex business processes. Method: This paper proposes Conversation Routines (CR), a framework that formalizes task-oriented dialogue logic via natural-language specifications, pioneering the integration of structured business workflows directly into LLM prompts—thereby decoupling dialogue design from tool implementation. CR supports modular routine definition and composition, natural-language-driven workflow orchestration, and synergistically combines tool-augmented conversational agents (Tool-Augmented CAS) with prompt engineering. Contribution/Results: Evaluated on two proof-of-concept scenarios—train ticket booking and interactive fault diagnosis—CR enables domain experts to build high-fidelity, high-task-success-rate dialogues without coding. It significantly improves system interpretability, reusability, and cross-role collaboration efficiency.
Domain-specific chatbots suffer from ambiguous user intent, contextual fragmentation, and interaction disorganization during multi-turn interactions—such as conditional filtering, multi-option selection, and comparative operations—due to the absence of GUI-like “submit/reset” mechanisms. To address this, this work introduces, for the first time, a form-based Submit/Reset paradigm into conversational systems, explicitly modeling user confirmation behaviors and context-switching actions. Methodologically, we integrate formalized state tracking, fine-grained user action recognition, and chain-of-thought (CoT) reasoning, augmented by prompt engineering to enhance large language models’ capacity for structured dialogue state representation. Experiments in hotel booking and customer management domains demonstrate significant improvements: +28.6% in multi-turn task coherence, +32.1% in user satisfaction, and a reduction of 2.4 turns on average, indicating enhanced operational efficiency.
Natural language generation (NLG) models suffer from poor interpretability due to architectural complexity and massive parameter counts, hindering their deployment in high-stakes decision-making. To address this, we propose the first unified human-computer interaction (HCI) and visualization classification framework specifically for NLG, systematically delineating three research paradigms and six core tasks that jointly address interpretability, controllability, and collaborative capability. By integrating principles from HCI design, eXplainable AI (XAI), information visualization, cognitive modeling, and NLG evaluation, we establish an interdisciplinary analytical framework that clarifies limitations of existing approaches. Furthermore, we introduce— for the first time in the large language model era—the “interactive–visualization co-enhancement” pathway. Our work provides both theoretical foundations and practical guidelines for developing trustworthy, controllable, and collaboratively capable NLG systems.
Autonomous driving testing requires efficient generation of simulation scenario code, yet non-programming domain experts struggle with this task. Method: We propose the first dialogue-based code generation system tailored for driving scenario modeling, integrating lightweight instruction fine-tuning, multi-turn dialogue state tracking, and domain-specific syntactic constraints to reliably translate natural language into symbolic scenario programs. Contribution/Results: Our work is the first to empirically validate the critical role of interactive dialogue in synthesizing complex scenario code. Leveraging minimal labeled data, the system achieves high-precision program synthesis. Human-in-the-loop experiments demonstrate that dialogue-based generation improves success rate by 4.5× over single-turn generation, significantly enhancing modeling accuracy, controllability, and practical usability for autonomous driving scenario creation.
This study addresses the challenge of assessing interaction quality in English-as-a-Second-Language (ESL) spoken dialogues. We propose the first interpretable evaluation framework that jointly models macro-level interaction labels (e.g., topic management) and micro-level linguistic features (e.g., pronouns, echoes, referring expressions). Leveraging manually annotated ESL dialogue data, we employ XGBoost and SVM to conduct feature importance analysis and build regression/classification models for interaction quality prediction. Our analysis systematically reveals, for the first time, how fine-grained linguistic signals predict high-level interaction quality: among 17 features—including reference words—several exhibit statistically significant effects, with pronoun usage demonstrating particularly strong predictive power. The findings empirically validate novel mappings between linguistic form and communicative competence, and enable automated, multidimensional, and interpretable assessment of ESL spoken interaction ability.
Existing large language models inadequately simulate authentic multi-turn human user behaviors—such as informal phrasing, personalized stylistic expression, and real-time self-correction—leading to biased and unrealistic evaluations of assistant models. Method: We propose User Language Models (User LMs), trained via human-centric post-training on multi-turn dialogue data—not by naïvely inverting assistant models—to explicitly capture realistic user interaction patterns. Contribution/Results: User LMs significantly improve behavioral fidelity and evaluation robustness, validated through both automated metrics and human assessments. Experiments reveal that when evaluated using User LMs, GPT-4o’s accuracy drops from 74.6% to 57.4% on programming and mathematical reasoning tasks, uncovering previously masked interaction weaknesses. This work establishes the first systematic, trustworthy user-side simulation paradigm for dialogue evaluation, introducing a new benchmark for assessing conversational capabilities of large models.
Current immersive analytics systems lack systematic characterization of embodied speech cues—such as spatial deixis, action verbs, and body-pose–linked terms—and how their dynamic interplay with disembodied language affects natural language interaction (NLI). Method: We conducted a Wizard-of-Oz study (N=24), collecting and axial-coding 1,280 speech acts to identify recurrent patterns. Contribution/Results: We propose the first taxonomy of five embodied–disembodied speech input patterns. Introducing semantic entropy as a metric, we quantify input uncertainty and reveal users’ adaptive switching among patterns based on task phase and embodiment dependency. Embodied cues significantly reduce semantic ambiguity and enhance system robustness in intent recognition. Our work establishes actionable design principles and an evaluation framework for embodied voice interaction in immersive analytics.
本文提出S^3-Bench,一个针对科学领域语音交互模型的评估框架,通过知识问答和多轮对话系统地解决专业术语理解和准确响应生成的问题。
This work addresses the challenge that existing testing methodologies struggle to effectively validate the behavior of conversational AI systems under interactive and multi-granularity integration scenarios. To this end, it proposes the first hierarchical testing framework specifically designed for conversational AI, systematically covering verification across multiple levels—from language and AI components, through single-agent systems, to multi-agent configurations. The approach integrates principles from software testing theory, architectural analysis of dialogue systems, and AI behavior validation techniques to construct a layered testing strategy. Experimental results demonstrate that the framework significantly enhances system reliability and consistency across different integration layers, offering a scalable and structured testing paradigm for complex conversational AI systems.
Traditional chat-based natural language interaction struggles to effectively handle structured information and complex tasks due to mismatches in data modality, high input entropy, and the absence of persistent state. This work proposes a novel paradigm—Software as Content (SaC)—which, for the first time, positions dynamically generated agent application interfaces as the central medium for human-agent interaction. By leveraging actionable interface elements to guide agent behavior and establishing a persistently evolvable shared interaction layer, SaC significantly enhances interaction efficiency and task adaptability across selection, exploration, and execution tasks. Grounded in a human-agent-environment interaction model and informed by interface evolution design principles, the approach integrates structured rendering with user feedback-driven mechanisms, thereby clarifying the operational boundaries within which natural language interaction excels.