🤖 AI Summary
This study addresses true full-duplex (TFD) speech interaction, aiming to achieve human-level simultaneous listening and speaking capabilities—enabling natural turn-taking, speech overlap, and real-time interruption. It tackles three key challenges in large-model-era full-duplex spoken language modeling: scarcity of synchronized multimodal data, architectural fragmentation, and disjointed evaluation criteria.
Method: We propose an “Engineered vs. Learned” dual-track taxonomy for synchronous modeling and establish the first unified, multidimensional evaluation framework encompassing temporal dynamics, behavioral arbitration mechanisms, semantic coherence, and acoustic fidelity. Our analysis systematically compares modular versus end-to-end architectures, conducts cross-model empirical studies, and models synchronized dialogue data.
Contribution/Results: The work delivers a reproducible benchmark for TFD speech systems and delineates a clear, principled technical roadmap—bridging gaps between architecture design, data curation, and holistic assessment in synchronous spoken interaction.
📝 Abstract
True Full-Duplex (TFD) voice communication--enabling simultaneous listening and speaking with natural turn-taking, overlapping speech, and interruptions--represents a critical milestone toward human-like AI interaction. This survey comprehensively reviews Full-Duplex Spoken Language Models (FD-SLMs) in the LLM era. We establish a taxonomy distinguishing Engineered Synchronization (modular architectures) from Learned Synchronization (end-to-end architectures), and unify fragmented evaluation approaches into a framework encompassing Temporal Dynamics, Behavioral Arbitration, Semantic Coherence, and Acoustic Performance. Through comparative analysis of mainstream FD-SLMs, we identify fundamental challenges: synchronous data scarcity, architectural divergence, and evaluation gaps, providing a roadmap for advancing human-AI communication.