π€ AI Summary
This study addresses the lack of adaptive mechanisms in existing full-duplex speech dialogue models, which struggle to meet dynamic interaction demands. Building upon Qwen3-Omni, this work proposes a three-tier collaborative adaptive full-duplex architecture. Methodologically, it introduces a compact token protocol and training-free runtime control to support dynamic window prediction and non-blocking inference. A three-stage progressive curriculum learning strategy is designed to decouple behavioral decision-making. Furthermore, the framework integrates dual-stream alignment, logits bias, GRPO reinforcement learning, and in-flight external reasoning techniques. Experimental results demonstrate that the proposed method surpasses state-of-the-art performance across multiple metrics on Full-Duplex-Bench v1 and v1.5, achieving a leading score of 72.9 on HumDial-FDBench.
π Abstract
Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which upgrades Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a further increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni and MiniCPM-o 4.5 on the majority of comparable turn-taking, overlap-behavior, and timing metrics, with gains in both interaction decisions and response timing. On the human-recorded HumDial-FDBench, it attains the highest Final score (72.9) of the compared duplex models.