multimodal dialogue simulation

Designs, builds, and evaluates systems and datasets that generate simulated conversations across multiple modalities (e.g., text, audio, transcripts) and associated annotations, producing paired audio/transcript data, speaker labels, timestamps, and other metadata. Creates multi-fidelity, diverse dialogue corpora by varying personas, accents, interaction roles and stages, and annotation granularity for use in training, testing, and analysis of conversational technologies.

multimodaldialoguesimulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Framework for Synthetic Audio Conversations Generation Using Large Language Models

Sep 02, 2024
KM
Kaung Myat Kyaw
🏛️ King Mongkut’s University of Technology Thonburi

Real conversational speech data for multi-speaker tasks—such as audio tagging, classification, and speaker identification—is scarce and expensive to annotate. Method: This paper proposes ConversaSynth, the first framework integrating multi-role large language model (LLM)-driven dialogue generation with high-fidelity text-to-speech (TTS) synthesis. It employs multi-role prompting, structured dialogue control, and topic-diverse sampling to generate end-to-end synthetic dialogue audio that is semantically coherent, speaker-discriminative, and acoustically natural. Contribution/Results: ConversaSynth jointly optimizes semantic consistency, speaker distinguishability, and speech naturalness. Experiments show that models trained on ConversaSynth-generated datasets achieve significantly improved downstream performance, with synthetic data approaching real conversations in both diversity and perceptual realism—establishing a new paradigm for high-quality synthetic data generation in low-resource multi-speaker speech tasks.

Creating diverse datasets for multi-speaker speech recognitionEnhancing audio tagging and classification modelsGenerating synthetic audio conversations using LLMs

Current speech dialogue systems suffer degraded performance in realistic, acoustically complex scenarios—such as audio mixing, background music interference, and emotional variability—primarily due to the scarcity of high-quality, multi-scenario conversational data. To address this, we introduce ShareChatX, the first large-scale synthetic speech dialogue dataset covering diverse acoustic conditions, and propose OmniChat, a unified dialogue system. Our approach features: (1) a novel synthetic-data-driven paradigm explicitly designed for complex acoustic environments; (2) a heterogeneous fusion module with dynamic feature selection that jointly models speech, musical context, and emotional states; and (3) an optimized training strategy integrating synthetic and real-world data. Evaluated on the real-world DailyTalk benchmark, OmniChat achieves state-of-the-art performance, demonstrating substantial improvements in audio event recognition, music-aware contextual understanding, and emotion expression modeling.

Complex Real-life ScenariosInsufficient Training DataSpeech Dialogue Systems

PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs

May 20, 2025
SI
Sho Inoue
🏛️ SRIBD | The Chinese University of Hong Kong | Nanjing University | National University of Singapore

The absence of personality annotations hinders personalized adaptation in voice-based dialogue systems. Method: This paper proposes the first personality-aware modeling framework for full-duplex voice dialogue. It integrates textual, acoustic, and behavioral multimodal cues and introduces an ASR- and LLM-driven end-to-end pipeline for automatic personality annotation generation. To enhance temporal coherence, it incorporates sequential modeling of emotion and response types, coupled with a human-in-the-loop annotation protocol for fine-grained personality trait prediction. Contribution/Results: The work establishes the first multimodal collaborative personality annotation framework, overcoming the key bottleneck of implicit-label-free automatic personality modeling. Human evaluation demonstrates significantly higher inter-annotator agreement compared to state-of-the-art baselines, empirically validating strong alignment between predicted personality traits and actual conversational behaviors.

Creating annotated datasets for personality-aware conversational agentsImproving alignment with human judgments in personality predictionPredicting personality traits from speech dialogues using multimodal cues

Existing conversational systems predominantly focus on text generation, neglecting prosodic expressivity and naturalness in speech output. This work addresses this gap by proposing the first human-like multimodal conversational agent designed for emotionally expressive speech responses. Methodologically, we (1) introduce the first multisensory dialogue dataset integrating linguistic, visual, and acoustic cues; (2) establish a novel speech generation paradigm that jointly models dialogue emotion and response style; and (3) leverage a multimodal large language model to generate text responses enriched with paralinguistic descriptions—explicitly encoding intonation, rhythm, and affect—which drive end-to-end speech synthesis. Experimental results demonstrate that audiovisual modality synergy significantly improves emotional fidelity and naturalness of synthesized speech. User studies confirm superior anthropomorphism and engagement compared to conventional TTS approaches.

Addressing lack of paralinguistic information in text-based responsesGenerating natural and engaging speech for conversational agentsIntegrating mood and style cues into multimodal speech generation

Latest Papers

What's happening recently
View more

This work proposes an unsupervised synthetic dialogue generation framework tailored for industrial settings where human-annotated data are scarce, relying solely on intent definitions. To enhance diversity, the approach explicitly incorporates topic and stylistic attributes and introduces two novel post-processing stylization models—Univ and Exam—combined with a large language model–based discriminative filtering mechanism to improve data quality. The study reveals that stylistic diversity has a significantly greater impact on the utility of synthetic data than topic diversity, and that integrating stylistic attributes during generation outperforms post-hoc style transfer. Experimental results demonstrate that the proposed method achieves 93.3% of the performance of models trained on human-annotated data across both industrial and public benchmarks, substantially enhancing the practicality of unlabeled synthetic dialogues.

annotation-freedialogue generationintent classification

This work addresses the high annotation cost and poor inter-annotator consistency in clinical physician–patient dialogue datasets, which hinder the evaluation of AI-based communication coding systems. To overcome these limitations, the authors propose a controllable generation framework for simulating clinician–patient dialogues with embedded behavioral annotations. The framework leverages predefined clinical scenarios, role-specific characteristics, and target communication behaviors, guided by dual codebooks—Global and WISER—and integrates speech synthesis with automatic audio quality assessment metrics (UTMOS, WV-MOS, WER, CER) and CLAP-based text–audio alignment. This approach enables, for the first time, multi-fidelity, interpretable, and reproducible clinical dialogue simulation. The system generates 3,388 cross-specialty dialogues, which automatic and human evaluations confirm exhibit high naturalness, transcription accuracy, and clinical authenticity, while also exposing limited sensitivity of existing coding systems along certain behavioral dimensions.

behavioral annotationclinical dialogue simulationcommunication coding

This work addresses the limitations in current AudioLLM development stemming from a scarcity of diverse, character-consistent, and instruction-aligned speech-text data, particularly regarding dialect coverage and speaker identity preservation. To overcome this, the authors propose a controllable generation framework that integrates World Values Survey–based persona construction, fine-grained dialogue scenario classification, and reference-audio-conditioned speech synthesis. Leveraging large language models, the framework generates multi-turn dialogues with consistent character traits and synthesizes speech conditioned on reference utterances to retain speaker characteristics and dialectal diversity. The project introduces MENASpeechBank, comprising 18,000 real utterances from 124 speakers across the Middle East and North Africa, alongside 417,000 high-quality synthetic dialogues spanning English, Modern Standard Arabic, and regional dialects. Evaluations confirm the data’s effectiveness, and all resources will be publicly released to advance community research.

AudioLLMsdialectal coveragemulti-speaker recordings

This work addresses the scarcity of training data and evaluation benchmarks for long-context audio reasoning, which hinders open-ended long-form audio generation and summarization. The authors propose the first end-to-end, open-source framework that synthesizes triadic medical consultations—comprising patient–clinician dialogues, multi-speaker audio, and structured clinical notes. The pipeline employs a role-playing large language model to generate initial-visit dialogues, which are then rendered into realistic multi-speaker speech incorporating overlapping utterances, pauses, room acoustics, and ambient noise, followed by automatic generation of SOAP-format clinical summaries. The project releases 8,800 synthetic dialogues (totaling 1,300 hours of audio) with corresponding reference summaries, filling a critical gap in medical long-audio datasets. Evaluations demonstrate that a cascaded approach significantly outperforms end-to-end models.

audio summarizationautomatic evaluationdoctor-patient conversations

Hot Scholars

YG

Yeyun Gong

Microsoft Research Asia
Natural Language GenerationQuestion AnsweringPre-training
GT

Gokhan Tur

University of Illinois Urbana-Champaign
Conversational AILanguage UnderstandingLarge Language Models
ZY

Zonghai Yao

Umass Amherst
Medical-LLMMulti-agent AI HospitalClinical ReasoningSynthetic Data
EY

Emine Yilmaz

University College London
Information RetrievalNatural Language ProcessingMachine Learning
LL

Le Li

Cornell University
Machine learningQuantitative researchStatistical modelingBioinformatics