Score
Designs, builds, and evaluates systems and datasets that generate simulated conversations across multiple modalities (e.g., text, audio, transcripts) and associated annotations, producing paired audio/transcript data, speaker labels, timestamps, and other metadata. Creates multi-fidelity, diverse dialogue corpora by varying personas, accents, interaction roles and stages, and annotation granularity for use in training, testing, and analysis of conversational technologies.
This work addresses the critical role of conversational user simulation in human-computer interaction, noting the absence of a systematic synthesis in existing literature. To bridge this gap, the paper proposes a unified classification framework grounded in large language models, integrating prior research along two key dimensions: user granularity and simulation objective. By structuring a comprehensive review around these axes, the study systematically examines core technical approaches, evaluation methodologies, and application scenarios. This structured analysis clarifies prevailing challenges and outlines promising future directions, thereby establishing a coherent research trajectory and theoretical foundation for advancing the field of conversational user simulation.
Real conversational speech data for multi-speaker tasks—such as audio tagging, classification, and speaker identification—is scarce and expensive to annotate. Method: This paper proposes ConversaSynth, the first framework integrating multi-role large language model (LLM)-driven dialogue generation with high-fidelity text-to-speech (TTS) synthesis. It employs multi-role prompting, structured dialogue control, and topic-diverse sampling to generate end-to-end synthetic dialogue audio that is semantically coherent, speaker-discriminative, and acoustically natural. Contribution/Results: ConversaSynth jointly optimizes semantic consistency, speaker distinguishability, and speech naturalness. Experiments show that models trained on ConversaSynth-generated datasets achieve significantly improved downstream performance, with synthetic data approaching real conversations in both diversity and perceptual realism—establishing a new paradigm for high-quality synthetic data generation in low-resource multi-speaker speech tasks.
Current speech dialogue systems suffer degraded performance in realistic, acoustically complex scenarios—such as audio mixing, background music interference, and emotional variability—primarily due to the scarcity of high-quality, multi-scenario conversational data. To address this, we introduce ShareChatX, the first large-scale synthetic speech dialogue dataset covering diverse acoustic conditions, and propose OmniChat, a unified dialogue system. Our approach features: (1) a novel synthetic-data-driven paradigm explicitly designed for complex acoustic environments; (2) a heterogeneous fusion module with dynamic feature selection that jointly models speech, musical context, and emotional states; and (3) an optimized training strategy integrating synthetic and real-world data. Evaluated on the real-world DailyTalk benchmark, OmniChat achieves state-of-the-art performance, demonstrating substantial improvements in audio event recognition, music-aware contextual understanding, and emotion expression modeling.
The absence of personality annotations hinders personalized adaptation in voice-based dialogue systems. Method: This paper proposes the first personality-aware modeling framework for full-duplex voice dialogue. It integrates textual, acoustic, and behavioral multimodal cues and introduces an ASR- and LLM-driven end-to-end pipeline for automatic personality annotation generation. To enhance temporal coherence, it incorporates sequential modeling of emotion and response types, coupled with a human-in-the-loop annotation protocol for fine-grained personality trait prediction. Contribution/Results: The work establishes the first multimodal collaborative personality annotation framework, overcoming the key bottleneck of implicit-label-free automatic personality modeling. Human evaluation demonstrates significantly higher inter-annotator agreement compared to state-of-the-art baselines, empirically validating strong alignment between predicted personality traits and actual conversational behaviors.
Existing conversational systems predominantly focus on text generation, neglecting prosodic expressivity and naturalness in speech output. This work addresses this gap by proposing the first human-like multimodal conversational agent designed for emotionally expressive speech responses. Methodologically, we (1) introduce the first multisensory dialogue dataset integrating linguistic, visual, and acoustic cues; (2) establish a novel speech generation paradigm that jointly models dialogue emotion and response style; and (3) leverage a multimodal large language model to generate text responses enriched with paralinguistic descriptions—explicitly encoding intonation, rhythm, and affect—which drive end-to-end speech synthesis. Experimental results demonstrate that audiovisual modality synergy significantly improves emotional fidelity and naturalness of synthesized speech. User studies confirm superior anthropomorphism and engagement compared to conventional TTS approaches.
This work proposes an unsupervised synthetic dialogue generation framework tailored for industrial settings where human-annotated data are scarce, relying solely on intent definitions. To enhance diversity, the approach explicitly incorporates topic and stylistic attributes and introduces two novel post-processing stylization models—Univ and Exam—combined with a large language model–based discriminative filtering mechanism to improve data quality. The study reveals that stylistic diversity has a significantly greater impact on the utility of synthetic data than topic diversity, and that integrating stylistic attributes during generation outperforms post-hoc style transfer. Experimental results demonstrate that the proposed method achieves 93.3% of the performance of models trained on human-annotated data across both industrial and public benchmarks, substantially enhancing the practicality of unlabeled synthetic dialogues.
This work addresses the high annotation cost and poor inter-annotator consistency in clinical physician–patient dialogue datasets, which hinder the evaluation of AI-based communication coding systems. To overcome these limitations, the authors propose a controllable generation framework for simulating clinician–patient dialogues with embedded behavioral annotations. The framework leverages predefined clinical scenarios, role-specific characteristics, and target communication behaviors, guided by dual codebooks—Global and WISER—and integrates speech synthesis with automatic audio quality assessment metrics (UTMOS, WV-MOS, WER, CER) and CLAP-based text–audio alignment. This approach enables, for the first time, multi-fidelity, interpretable, and reproducible clinical dialogue simulation. The system generates 3,388 cross-specialty dialogues, which automatic and human evaluations confirm exhibit high naturalness, transcription accuracy, and clinical authenticity, while also exposing limited sensitivity of existing coding systems along certain behavioral dimensions.
This work addresses the limitations in current AudioLLM development stemming from a scarcity of diverse, character-consistent, and instruction-aligned speech-text data, particularly regarding dialect coverage and speaker identity preservation. To overcome this, the authors propose a controllable generation framework that integrates World Values Survey–based persona construction, fine-grained dialogue scenario classification, and reference-audio-conditioned speech synthesis. Leveraging large language models, the framework generates multi-turn dialogues with consistent character traits and synthesizes speech conditioned on reference utterances to retain speaker characteristics and dialectal diversity. The project introduces MENASpeechBank, comprising 18,000 real utterances from 124 speakers across the Middle East and North Africa, alongside 417,000 high-quality synthetic dialogues spanning English, Modern Standard Arabic, and regional dialects. Evaluations confirm the data’s effectiveness, and all resources will be publicly released to advance community research.
This work addresses the scarcity of training data and evaluation benchmarks for long-context audio reasoning, which hinders open-ended long-form audio generation and summarization. The authors propose the first end-to-end, open-source framework that synthesizes triadic medical consultations—comprising patient–clinician dialogues, multi-speaker audio, and structured clinical notes. The pipeline employs a role-playing large language model to generate initial-visit dialogues, which are then rendered into realistic multi-speaker speech incorporating overlapping utterances, pauses, room acoustics, and ambient noise, followed by automatic generation of SOAP-format clinical summaries. The project releases 8,800 synthetic dialogues (totaling 1,300 hours of audio) with corresponding reference summaries, filling a critical gap in medical long-audio datasets. Evaluations demonstrate that a cascaded approach significantly outperforms end-to-end models.