Score
Designs and implements conversational interfaces and supporting back-end components that accept and parse natural-language questions about model behavior, infer user intent, and manage an explainability dialogue. Produces concise human-readable explanation responses and coordinates intent-aligned visualizations or follow-up prompts so users can iteratively probe and understand model decisions.
This paper addresses the lack of systematic design and evaluation frameworks for prompt-based natural language explanations (NLEs) in transparent AI governance. We propose the first taxonomy specifically tailored to prompt-driven NLEs, structured along three dimensions: context, generation & presentation, and evaluation. The taxonomy integrates task objectives, data characteristics, stakeholder needs, generation mechanisms, interaction modalities, output formats, and user-centered evaluation criteria into a multidimensional, structured classification. Its key innovation lies in adapting eXplainable AI (XAI) taxonomic principles to the prompt-based NLE paradigm—moving beyond conventional model-internal explanations. The resulting framework provides researchers, auditors, and policymakers with actionable design guidelines and standardized evaluation benchmarks, thereby enhancing the practical utility of NLEs in transparency, auditability, and regulatory compliance. (149 words)
This study addresses the human-centered explainability challenge in interactive information systems, aiming to enhance users’ comprehension, interpretation, and critical evaluation of AI outputs to support informed decision-making. Following the PRISMA guidelines, we systematically screened and structurally coded 100 peer-reviewed articles. Based on this analysis, we propose a novel five-dimensional conceptual model of explainability—comprising transparency, causality, relevance, controllability, and actionability—and introduce the first user-needs-driven taxonomy for explanation design. Furthermore, we identify six user-centered dimensions for explainability evaluation—the first such classification in the literature. These contributions advance explainability research from ad hoc, experience-based practice toward systematic, theory-grounded inquiry. The resulting integrative framework provides both theoretical rigor and practical guidance for designing transparent, trustworthy, and responsible interactive information systems.
Domain-specific chatbots suffer from ambiguous user intent, contextual fragmentation, and interaction disorganization during multi-turn interactions—such as conditional filtering, multi-option selection, and comparative operations—due to the absence of GUI-like “submit/reset” mechanisms. To address this, this work introduces, for the first time, a form-based Submit/Reset paradigm into conversational systems, explicitly modeling user confirmation behaviors and context-switching actions. Methodologically, we integrate formalized state tracking, fine-grained user action recognition, and chain-of-thought (CoT) reasoning, augmented by prompt engineering to enhance large language models’ capacity for structured dialogue state representation. Experiments in hotel booking and customer management domains demonstrate significant improvements: +28.6% in multi-turn task coherence, +32.1% in user satisfaction, and a reduction of 2.4 turns on average, indicating enhanced operational efficiency.
Large language models (LLMs) exhibit unstable dialogue behavior and poor maintainability in complex business processes. Method: This paper proposes Conversation Routines (CR), a framework that formalizes task-oriented dialogue logic via natural-language specifications, pioneering the integration of structured business workflows directly into LLM prompts—thereby decoupling dialogue design from tool implementation. CR supports modular routine definition and composition, natural-language-driven workflow orchestration, and synergistically combines tool-augmented conversational agents (Tool-Augmented CAS) with prompt engineering. Contribution/Results: Evaluated on two proof-of-concept scenarios—train ticket booking and interactive fault diagnosis—CR enables domain experts to build high-fidelity, high-task-success-rate dialogues without coding. It significantly improves system interpretability, reusability, and cross-role collaboration efficiency.
This study investigates whether explicit intent recognition is a necessary prerequisite for generating high-quality responses in task-oriented dialogue systems. Method: Challenging the conventional intent-module-dependent paradigm, we propose and comparatively evaluate two strategies—“intent-first” and “end-to-end direct generation”—using large language models (e.g., T5) fine-tuned on multiple public task-oriented dialogue datasets. Evaluation encompasses both linguistic quality and task completion rate. Contribution/Results: Experiments demonstrate that, in typical service scenarios, direct generation achieves performance on par with or exceeding that of the intent-first approach—even without intent annotations—while substantially reducing system complexity and inference latency. These findings empirically challenge the assumed necessity of explicit intent recognition, providing evidence and conceptual support for lightweight, low-latency service assistant design grounded in a new end-to-end paradigm.
This work addresses a critical limitation in current explainable artificial intelligence (XAI) research, which predominantly treats explanations as static outputs while overlooking their inherently dynamic and dialogic nature in real-world settings. To bridge this gap, the paper introduces a Human-Centered Conversational XAI (HC²XAI) framework that systematically integrates dialogue as a core dimension of XAI for the first time. Emphasizing interactivity, context dependence, and contextual awareness, HC²XAI synthesizes principles from conversational system design, human-computer interaction theory, and established XAI techniques to enable multi-turn, adaptive explanation generation. The framework’s efficacy is demonstrated through three representative application scenarios, offering a novel paradigm that shifts XAI from static explanations toward dynamic, user-responsive dialogue and opening new avenues for research within the conversational user interface (CUI) community.
本文提出一种三代理框架,通过模拟对话评估和改进大型语言模型在处理模糊问题时的澄清能力,以提高交互系统的用户意图理解准确性。
This study addresses the cumbersome interaction and difficulty of exploring complex systems via natural language in 3D software visualization. To overcome these limitations, this work proposes integrating a large language model (LLM)-based intelligent assistant into ExplorViz. Drawing upon the Model Context Protocol (MCP) paradigm, the approach combines probabilistic LLM reasoning with deterministic tool invocation, enabling users to query software architectures in natural language and trigger visualization operations such as highlighting and view adjustment for real-time feedback-driven interactive exploration. Experimental results demonstrate that the assistant performs reliably in structural summarization and targeted operations, significantly reducing the interaction overhead associated with program comprehension.
This work addresses a key challenge in automated planning: enabling AI systems to provide human planners with intelligible and trustworthy explanations through natural dialogue, thereby supporting preference- and expertise-driven guidance. We propose the first conversational, context-aware multi-agent large language model (LLM) framework capable of dynamically generating personalized explanations without relying on predefined templates, while effectively responding to user interactions. By deeply integrating planning systems with natural language interaction—particularly instantiated for scenarios involving goal conflicts—our approach facilitates more intuitive human–AI collaboration. User studies demonstrate that, compared to conventional template-based explanation interfaces, our method significantly enhances users’ understanding of proposed plans and their trust in the system.
This study addresses the limited effectiveness of existing explainable artificial intelligence (XAI) approaches in improving users’ objective performance and the lack of empirical evidence on how conversational XAI influences prediction accuracy, model understanding, and error identification. Through a controlled between-subjects experiment employing intrinsically interpretable models, the authors evaluate the impact of conversational versus question-answering XAI on user decision-making, enabling participants to detect and correct systematic model errors. Preliminary results (N=42) indicate that both XAI modalities significantly enhance user performance compared to baseline conditions, though no significant difference emerges between them; notably, user engagement remains low, suggesting directions for refined intervention strategies. This work contributes an experimental paradigm that empowers users to surpass model performance and provides empirical validation of conversational XAI’s efficacy.
This work proposes a multimodal learning framework based on adaptive context fusion to address the limited generalization of existing methods in complex scenarios. The approach dynamically aligns visual and linguistic features and incorporates a lightweight gating mechanism to enable efficient cross-modal integration. Experimental results demonstrate that the model significantly outperforms current state-of-the-art methods across multiple benchmark datasets, achieving improvements of 3.2% in accuracy and 5.7% in robustness. The primary contribution lies in the design of a scalable fusion architecture that effectively mitigates the semantic gap between modalities, offering a novel technical pathway for multimodal understanding tasks.
This study addresses the challenge of fine-grained investigation into dynamic user interactions with large language models (LLMs) in realistic yet controlled settings, particularly the lack of joint control over prompt construction behaviors and model responses. To bridge this gap, the authors present a configurable experimental platform that enables precise manipulation of LLM behavioral parameters while simultaneously capturing multimodal data—including keystroke-level prompt editing actions (e.g., typing, deletions, pauses), interaction timing, and subjective user feedback. The platform uniquely supports coordinated control over model behavior, user input, and contextual conditions during LLM interactions, integrating behavioral logs with self-report measures to facilitate causal inference in AI interface design. Empirical validation demonstrates that technical detail level and response length significantly affect user memory retention, confirming the platform’s efficacy for systematically evaluating design variables in AI-human interaction.