AdaptDuplex: from static to adaptive full-duplex spoken dialogue
This study addresses the lack of adaptive mechanisms in existing full-duplex speech dialogue models, which struggle to meet dynamic interaction demands. Building upon Qwen3-Omni, this work proposes a three-tier collaborative adaptive full-duplex architecture. Methodologically, it introduces a compact token protocol and training-free runtime control to support dynamic window prediction and non-blocking inference. A three-stage progressive curriculum learning strategy is designed to decouple behavioral decision-making. Furthermore, the framework integrates dual-stream alignment, logits bias, GRPO reinforcement learning, and in-flight external reasoning techniques. Experimental results demonstrate that the proposed method surpasses state-of-the-art performance across multiple metrics on Full-Duplex-Bench v1 and v1.5, achieving a leading score of 72.9 on HumDial-FDBench.