From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models

📅 2025-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses true full-duplex (TFD) speech interaction, aiming to achieve human-level simultaneous listening and speaking capabilities—enabling natural turn-taking, speech overlap, and real-time interruption. It tackles three key challenges in large-model-era full-duplex spoken language modeling: scarcity of synchronized multimodal data, architectural fragmentation, and disjointed evaluation criteria. Method: We propose an “Engineered vs. Learned” dual-track taxonomy for synchronous modeling and establish the first unified, multidimensional evaluation framework encompassing temporal dynamics, behavioral arbitration mechanisms, semantic coherence, and acoustic fidelity. Our analysis systematically compares modular versus end-to-end architectures, conducts cross-model empirical studies, and models synchronized dialogue data. Contribution/Results: The work delivers a reproducible benchmark for TFD speech systems and delineates a clear, principled technical roadmap—bridging gaps between architecture design, data curation, and holistic assessment in synchronous spoken interaction.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
True Full-Duplex (TFD) voice communication--enabling simultaneous listening and speaking with natural turn-taking, overlapping speech, and interruptions--represents a critical milestone toward human-like AI interaction. This survey comprehensively reviews Full-Duplex Spoken Language Models (FD-SLMs) in the LLM era. We establish a taxonomy distinguishing Engineered Synchronization (modular architectures) from Learned Synchronization (end-to-end architectures), and unify fragmented evaluation approaches into a framework encompassing Temporal Dynamics, Behavioral Arbitration, Semantic Coherence, and Acoustic Performance. Through comparative analysis of mainstream FD-SLMs, we identify fundamental challenges: synchronous data scarcity, architectural divergence, and evaluation gaps, providing a roadmap for advancing human-AI communication.
Problem

Research questions and friction points this paper is trying to address.

Achieving true full-duplex voice communication with simultaneous listening and speaking
Addressing synchronous data scarcity and architectural divergence challenges
Unifying fragmented evaluation approaches for spoken language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Engineered and Learned Synchronization architectures
Unified evaluation framework with four metrics
Addressing data scarcity and architectural divergence challenges
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yuxuan Chen
Jilin University, Changchun, China
H
Haoyuan Yu
Hunan University, Changsha, China