build full-duplex slms

Design and implement end-to-end spoken language models that operate in native full‑duplex (simultaneous bidirectional/streaming) settings, including models and frameworks that perform both continuous input capture and continuous output generation. Build and optimize full‑duplex inference pipelines and architectures to minimize latency, handle interruptions and turn‑taking, and support tasks such as spoken question answering and fluid interactive dialogue.

buildfull-duplexslms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing speech-language models predominantly operate in unidirectional, turn-taking paradigms, lacking real-time interjection capability and synchronous response. This work introduces the first end-to-end duplex speech-to-speech (S2S) architecture, eliminating the need for pre-trained speech modules and directly modeling concurrent user and agent speech streams. Methodologically, it employs a streaming encoder, separate user/agent modeling, codec-channel fusion, and an LLM-driven duplex generation mechanism. Key contributions include: (1) the first purely end-to-end duplex S2S paradigm; (2) the first publicly released complete training and inference codebase; (3) high-fidelity speech synthesis at an ultra-low bitrate of 0.6 kbps; and (4) drastically reduced data requirements, enabling rapid adaptation to arbitrary LLMs. Experiments demonstrate substantial improvements over prior duplex approaches in interjection latency, turn-taking control accuracy, and speech naturalness.

Eliminates speech pretrain need, simplifying duplex S2S model developmentEnables real-time duplex speech interaction with barge-in capabilityReduces bitrate and improves agent voice quality via codec fine-tuning

This work addresses the scarcity of high-quality, speaker-separated full-duplex conversational speech data—a critical bottleneck in training spoken dialogue language models—given that most existing large-scale public speech corpora are monaural and lack explicit speaker turn structure. To bridge this gap, the authors introduce the DuplexChat project, which presents the first large-scale effort to construct speaker-separated, full-duplex conversational datasets from massive monaural podcast archives. They develop DuplexChat-Pipe, a comprehensive pipeline integrating language filtering, audio cleaning, diarization-guided two-speaker segment extraction, and speech separation with restoration. The resulting corpus comprises 282,634 hours of English and 132,723 hours of Japanese conversational speech, faithfully preserving natural turn-taking dynamics and substantially advancing resource availability for spoken dialogue research.

full-duplexlanguage modelingspeaker separation

Existing speech dialogue systems struggle to emulate natural human full-duplex interaction—such as interjections, speech overlaps, and immediate turn-taking. Method: This paper proposes an end-to-end full-duplex speech dialogue system that requires no architectural modification to the GPT backbone. It introduces a novel three-stage post-training paradigm: (i) cross-modal alignment, (ii) progressive learning from half-duplex to full-duplex behavior, and (iii) unified “flattening” of speech-text joint representations. The system performs speech-text joint modeling using a pure text-based large language model, enabling real-time bidirectional speech input and output. Contribution/Results: Experiments demonstrate significant reduction in end-to-end latency, improved speech naturalness and interaction fluency, and high-quality synchronous interaction—all while preserving model compatibility. Code and audio examples are publicly released.

Full-duplex conversationNatural language processingSpeech interaction

Latest Papers

What's happening recently
View more

This work addresses the challenge of maintaining semantic coherence in full-duplex spoken dialogue, where large language models struggle to generate consistent responses while simultaneously processing streaming user speech, often suffering from contextual interference. The study introduces user-stream routing as a foundational modeling dimension and presents a unified framework for full-duplex spoken dialogue, comparing two strategies: channel fusion—directly injecting the user stream into the generation process—and cross-attention routing—accessing external memory via adapter modules. Experimental results demonstrate that channel fusion achieves superior semantic understanding in spoken question answering but is highly sensitive to interruptions, whereas cross-attention routing, despite slightly lower task performance, substantially enhances response coherence and contextual robustness, revealing a critical trade-off between semantic integration capability and robustness in real-time conversational systems.

context corruptionfull-duplex spoken dialoguelarge language models

This work addresses a critical challenge in full-duplex spoken language modeling: the sharing of deep-layer parameters between acoustic and semantic modalities often induces gradient conflict, leading to knowledge degradation and compromised semantic integrity. The study is the first to uncover the underlying mechanism of this cross-modal interference and introduces Lychee-FD, a novel framework that decouples modalities through hierarchical parameter separation in deep layers while preserving cross-modal consistency via a dedicated semantic alignment channel. This enables native end-to-end full-duplex modeling without sacrificing coherence. Experimental results demonstrate that Lychee-FD achieves state-of-the-art performance across multiple benchmarks, yielding a 7.4% absolute improvement in spoken question-answering accuracy and a 28.5% gain in interaction fluency, all while maintaining efficient inference.

acoustic-semantic modelingfull-duplex Spoken Language Modelsmodality interference

This work addresses the scarcity of high-quality, multi-speaker natural conversation audio data—a key bottleneck in developing full-duplex speech language models. Existing datasets are often limited to single speakers or small scales, and standard preprocessing pipelines are prone to speaker diarization errors and ASR hallucinations. To overcome these challenges, we propose the first open-source, end-to-end scalable preprocessing framework tailored for full-duplex speech language modeling. Our approach integrates robust speaker separation, multi-channel speech alignment, hallucination-resistant ASR post-processing, and dialogue structure modeling. This pipeline substantially mitigates speaker confusion and recognition errors, enabling the generation of high-quality, multi-turn, multi-speaker conversational datasets. The resulting data provides a reliable foundation for training full-duplex models, significantly enhancing interaction naturalness and real-time responsiveness.

ASR hallucinationsdiarization errorsfull-duplex

Existing models struggle to simultaneously support real-time interaction, complex reasoning, and tool use within a full-duplex multimodal framework. This work proposes an asynchronous full-duplex architecture that decouples interaction from reasoning: an interaction layer processes audio and video inputs in an end-to-end streaming fashion to generate immediate responses, while a thinking layer employs plug-in modules to perform asynchronous, complex reasoning and tool invocation. To facilitate training, the authors introduce a Writer-Director pipeline for constructing continuous interactive data. Evaluated on multiple public benchmarks, the system demonstrates strong performance, significantly enhancing the naturalness and fluency of multimodal interactions.

complex reasoningcontinuous interactionfull-duplex interaction

Hot Scholars

JX

Jiacheng Xu

Nanyang Technological University
Reinforcement LearningLarge Language Model
WZ

Wenjiang Zhou

Peking University, HUST
AI for scienceAtomistic simulationsSuper-Planckian far-field
AO

Aydogan Ozcan

Chancellor's Professor at UCLA & HHMI Professor
Computational ImagingHolographyMicroscopySensing
HG

Heting Gao

University of Illinois at Urbana-Champaign
Speech Recognition
LL

Lijiang Li

Xiamen University
Diffusion Models