Score
Design and implement end-to-end spoken language models that operate in native full‑duplex (simultaneous bidirectional/streaming) settings, including models and frameworks that perform both continuous input capture and continuous output generation. Build and optimize full‑duplex inference pipelines and architectures to minimize latency, handle interruptions and turn‑taking, and support tasks such as spoken question answering and fluid interactive dialogue.
This study addresses true full-duplex (TFD) speech interaction, aiming to achieve human-level simultaneous listening and speaking capabilities—enabling natural turn-taking, speech overlap, and real-time interruption. It tackles three key challenges in large-model-era full-duplex spoken language modeling: scarcity of synchronized multimodal data, architectural fragmentation, and disjointed evaluation criteria. Method: We propose an “Engineered vs. Learned” dual-track taxonomy for synchronous modeling and establish the first unified, multidimensional evaluation framework encompassing temporal dynamics, behavioral arbitration mechanisms, semantic coherence, and acoustic fidelity. Our analysis systematically compares modular versus end-to-end architectures, conducts cross-model empirical studies, and models synchronized dialogue data. Contribution/Results: The work delivers a reproducible benchmark for TFD speech systems and delineates a clear, principled technical roadmap—bridging gaps between architecture design, data curation, and holistic assessment in synchronous spoken interaction.
Existing speech-language models predominantly operate in unidirectional, turn-taking paradigms, lacking real-time interjection capability and synchronous response. This work introduces the first end-to-end duplex speech-to-speech (S2S) architecture, eliminating the need for pre-trained speech modules and directly modeling concurrent user and agent speech streams. Methodologically, it employs a streaming encoder, separate user/agent modeling, codec-channel fusion, and an LLM-driven duplex generation mechanism. Key contributions include: (1) the first purely end-to-end duplex S2S paradigm; (2) the first publicly released complete training and inference codebase; (3) high-fidelity speech synthesis at an ultra-low bitrate of 0.6 kbps; and (4) drastically reduced data requirements, enabling rapid adaptation to arbitrary LLMs. Experiments demonstrate substantial improvements over prior duplex approaches in interjection latency, turn-taking control accuracy, and speech naturalness.
This work addresses the scarcity of high-quality, speaker-separated full-duplex conversational speech data—a critical bottleneck in training spoken dialogue language models—given that most existing large-scale public speech corpora are monaural and lack explicit speaker turn structure. To bridge this gap, the authors introduce the DuplexChat project, which presents the first large-scale effort to construct speaker-separated, full-duplex conversational datasets from massive monaural podcast archives. They develop DuplexChat-Pipe, a comprehensive pipeline integrating language filtering, audio cleaning, diarization-guided two-speaker segment extraction, and speech separation with restoration. The resulting corpus comprises 282,634 hours of English and 132,723 hours of Japanese conversational speech, faithfully preserving natural turn-taking dynamics and substantially advancing resource availability for spoken dialogue research.
Existing speech dialogue systems struggle to emulate natural human full-duplex interaction—such as interjections, speech overlaps, and immediate turn-taking. Method: This paper proposes an end-to-end full-duplex speech dialogue system that requires no architectural modification to the GPT backbone. It introduces a novel three-stage post-training paradigm: (i) cross-modal alignment, (ii) progressive learning from half-duplex to full-duplex behavior, and (iii) unified “flattening” of speech-text joint representations. The system performs speech-text joint modeling using a pure text-based large language model, enabling real-time bidirectional speech input and output. Contribution/Results: Experiments demonstrate significant reduction in end-to-end latency, improved speech naturalness and interaction fluency, and high-quality synchronous interaction—all while preserving model compatibility. Code and audio examples are publicly released.
This work addresses the challenge of maintaining semantic coherence in full-duplex spoken dialogue, where large language models struggle to generate consistent responses while simultaneously processing streaming user speech, often suffering from contextual interference. The study introduces user-stream routing as a foundational modeling dimension and presents a unified framework for full-duplex spoken dialogue, comparing two strategies: channel fusion—directly injecting the user stream into the generation process—and cross-attention routing—accessing external memory via adapter modules. Experimental results demonstrate that channel fusion achieves superior semantic understanding in spoken question answering but is highly sensitive to interruptions, whereas cross-attention routing, despite slightly lower task performance, substantially enhances response coherence and contextual robustness, revealing a critical trade-off between semantic integration capability and robustness in real-time conversational systems.
This work addresses a critical challenge in full-duplex spoken language modeling: the sharing of deep-layer parameters between acoustic and semantic modalities often induces gradient conflict, leading to knowledge degradation and compromised semantic integrity. The study is the first to uncover the underlying mechanism of this cross-modal interference and introduces Lychee-FD, a novel framework that decouples modalities through hierarchical parameter separation in deep layers while preserving cross-modal consistency via a dedicated semantic alignment channel. This enables native end-to-end full-duplex modeling without sacrificing coherence. Experimental results demonstrate that Lychee-FD achieves state-of-the-art performance across multiple benchmarks, yielding a 7.4% absolute improvement in spoken question-answering accuracy and a 28.5% gain in interaction fluency, all while maintaining efficient inference.
This work addresses the scarcity of high-quality, multi-speaker natural conversation audio data—a key bottleneck in developing full-duplex speech language models. Existing datasets are often limited to single speakers or small scales, and standard preprocessing pipelines are prone to speaker diarization errors and ASR hallucinations. To overcome these challenges, we propose the first open-source, end-to-end scalable preprocessing framework tailored for full-duplex speech language modeling. Our approach integrates robust speaker separation, multi-channel speech alignment, hallucination-resistant ASR post-processing, and dialogue structure modeling. This pipeline substantially mitigates speaker confusion and recognition errors, enabling the generation of high-quality, multi-turn, multi-speaker conversational datasets. The resulting data provides a reliable foundation for training full-duplex models, significantly enhancing interaction naturalness and real-time responsiveness.
Existing models struggle to simultaneously support real-time interaction, complex reasoning, and tool use within a full-duplex multimodal framework. This work proposes an asynchronous full-duplex architecture that decouples interaction from reasoning: an interaction layer processes audio and video inputs in an end-to-end streaming fashion to generate immediate responses, while a thinking layer employs plug-in modules to perform asynchronous, complex reasoning and tool invocation. To facilitate training, the authors introduce a Writer-Director pipeline for constructing continuous interactive data. Evaluated on multiple public benchmarks, the system demonstrates strong performance, significantly enhancing the naturalness and fluency of multimodal interactions.