Score
Designs and executes processes to collect, clean, annotate, and structure conversational datasets and to analyze multi-turn interaction patterns and model behavior. Work includes collecting multi-party interactions, segmenting dialogue into turns, normalizing speakers and timestamps, tokenizing and annotating dialog acts and conversational cues, filtering noise and irrelevant tokens, and extracting style or identity features for analysis or downstream systems.
Service workflows in customer service dialogues are often missing and unstructured, leading to inconsistent AI responses. Method: This paper proposes the first integrated framework for automatic dialogue workflow extraction and simulation-based evaluation. It combines retrieval-augmented generation (RAG) with a novel question-answering–based chain-of-thought (QA-CoT) prompting technique to improve structured workflow generation accuracy. Additionally, it introduces a scalable two-agent (agent + customer) simulation evaluation mechanism for automated, large-scale workflow assessment. Contribution/Results: We design a macro-accuracy metric that achieves high agreement with human evaluation (Spearman ρ > 0.92). On the ABCD and SynthABCD datasets, our method improves average macro-accuracy by 12.16%, significantly mitigating response inconsistency caused by workflow omission in service AI systems.
This study addresses the challenge of automatically analyzing collaborative processes in task-oriented human-human dialogues by systematically reviewing discourse-level approaches to collaboration analysis. Integrating methods from conversation analysis, collaboration theory, behavioral coding schemes, and computational modeling, the work presents the first multidimensional and structured knowledge framework that comprehensively organizes existing techniques in terms of their task formulations, modeling strategies, and respective strengths and limitations. By offering a clear research roadmap and practical reference for the field, this contribution not only clarifies current methodological constraints but also identifies promising directions for future inquiry, thereby advancing collaboration analysis toward greater systematicity and computational tractability.
This work addresses the limitations of linear conversation logs generated by existing conversational data analysis systems, which hinder data workers’ ability to retrospect and communicate about nonlinear, iterative analytical processes. To overcome this, the paper proposes a structured dialogue presentation method that introduces probes enabling multi-level navigation, on-demand detail expansion, and context-enhanced summarization—going beyond conventional scrolling and keyword search. By integrating visual recall with sequential and abstraction-based navigation strategies, the approach effectively supports users in recalling, reorienting within, and prioritizing past analytical exchanges. A user study with ten participants demonstrates that the method significantly enhances traceability of analytical reasoning and improves collaborative efficiency, validating its effectiveness in real-world data analysis workflows.
Large language models (LLMs) exhibit unstable dialogue behavior and poor maintainability in complex business processes. Method: This paper proposes Conversation Routines (CR), a framework that formalizes task-oriented dialogue logic via natural-language specifications, pioneering the integration of structured business workflows directly into LLM prompts—thereby decoupling dialogue design from tool implementation. CR supports modular routine definition and composition, natural-language-driven workflow orchestration, and synergistically combines tool-augmented conversational agents (Tool-Augmented CAS) with prompt engineering. Contribution/Results: Evaluated on two proof-of-concept scenarios—train ticket booking and interactive fault diagnosis—CR enables domain experts to build high-fidelity, high-task-success-rate dialogues without coding. It significantly improves system interpretability, reusability, and cross-role collaboration efficiency.
Existing dialogue datasets generally lack fine-grained annotations of multi-topic evolution and natural topic transitions, hindering dynamic topic identification and modeling in long conversations. To address this, we propose a controlled dialogue collection paradigm that explicitly supports multi-turn topic emergence and dynamic switching—novel in its design. Leveraging a custom-built instant messaging platform, a structured elicitation protocol, and an intent-aware conversational topic annotation scheme, we construct the first high-quality, topic-analyzed long-dialogue corpus. This corpus features explicit temporal structure, precisely annotated topic boundaries, and fine-grained topic transition types (e.g., shift, continuation, elaboration). It fills a critical empirical data gap in spoken dialogue topic structure research and establishes a robust foundation for topic identification, tracking, and computational modeling.
This work addresses the limited transparency and interpretability of conventional large language models (LLMs) in business process modeling, particularly their difficulty in handling complex dependencies. To overcome these limitations, the paper proposes a hybrid, interpretable process modeling paradigm that reframes model construction as an iterative dialogue between human experts and LLMs. This approach integrates task decomposition, intermediate artifact generation, and explicit documentation of modeling rationale, combined with behavioral relationship analysis and specialized modeling tools, to incrementally construct structured process models. Building on this paradigm, the authors develop Pragmos, a prototype system that enables users and LLMs to collaboratively create high-quality business process models that are transparent, interpretable, and capable of progressive evolution.
This study addresses core challenges in task-oriented dialogue with large language models (LLMs), including weak topical coherence, insufficient knowledge progression, inconsistent role embodiment, and coarse-grained controllability. To this end, we propose the first systematic, multi-dimensional parametric framework for dialogue quality control. The framework defines nine quantifiable and intervenable control parameters across six dimensions—semantic coherence, knowledge evolution, role consistency, among others—enabling fine-grained, reproducible modeling and regulation of dialogue attributes. Empirical evaluation on mainstream LLMs demonstrates statistically significant improvements in dialogue quality (p < 0.01) and task adaptability. The framework supports diverse application scenarios, including education, psychological counseling, customer service, and entertainment. By establishing a standardized, parameter-driven paradigm for dialogue generation quality control, this work advances controllable, reliable, and domain-adaptable conversational AI.
This study addresses the lack of systematic preprocessing standards, integrated analytical workflows, and cross-method consistency checks in current computer-based assessment process data. To bridge this gap, the authors propose an end-to-end analytical framework featuring a unified preprocessing pipeline and a dual-path analysis paradigm that synergistically combines feature engineering with model-based inference. The framework incorporates large language models (LLMs) to standardize action sequences and facilitate process-data-driven differential item functioning (DIF) detection. Technically, it integrates timestamp correction, action chunking, n-gram and TF-IDF feature extraction, multidimensional scaling, hidden Markov modeling, and subtask identification. Empirical results demonstrate that n-gram–based behavioral clustering offers diagnostic value for incorrect responders, multidimensional scaling effectively reconstructs behavioral constructs, and process data can identify and mitigate construct-irrelevant group differences.
This study investigates whether intelligent agents endowed with reflection and memory capabilities can achieve observable and controllable performance gains in information extraction tasks. Focusing on structured data extraction from academic paper PDFs, we propose an enhanced agent architecture, S2, which incorporates a dynamic tool selection mechanism and an expanded set of PDF processing tools. We further introduce an evaluation framework centered on behavioral controllability. Experimental results demonstrate that the proposed agent effectively adapts its execution strategy through reflection, retrying, and memory mechanisms, significantly outperforming fixed-pipeline large language model workflows on critical failure modes, thereby validating the efficacy and superiority of our design.
This paper addresses the limitations of existing topic detection methods for large-scale dialogues (e.g., customer service, sales), which rely on predefined intent taxonomies and lack user controllability. We propose a controllable topic detection framework that jointly clusters dialogue texts with large language model (LLM) representations, dynamically adjusts topic granularity via user-provided preference data, and supports flexible topic formulation and personalization. Evaluation employs a multi-dimensional metric combining automated scoring and human verification. Our key contributions are: (1) the first user-adjustable granularity paradigm for dialogue topic detection; (2) the release of the DSTC 12 Topic Detection task—including a real-world dialogue dataset, open-source system, and standardized evaluation benchmark; and (3) empirical validation across multiple participating teams, demonstrating significant improvements in detection accuracy, interpretability, and practical extensibility, while substantially reducing manual analysis effort.