🤖 AI Summary
Autoregressive inference in full-duplex spoken dialogue models incurs substantial computational overhead, hindering real-time interaction. To address this limitation, this work proposes a rolling mask diffusion framework that parallelizes the prediction of multiple future frames within a single backbone network invocation, thereby significantly reducing sequential computation costs. Furthermore, a dynamic verification and local correction mechanism is introduced to recompute only unplayed segments, effectively balancing low latency with generation quality. By integrating confident prefix consumption with a dual-strategy inference approach, the proposed method achieves up to 1.8× core acceleration while satisfying an 80ms real-time constraint, all without compromising conversational naturalness or dialogue quality.
📝 Abstract
Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-LISTEN consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-SPEAK can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-LISTEN and DiffuPlex-SPEAK achieve $1.46\times$ and $1.59\times$ deployment-path wall-clock speedups and $1.61\times$ and $1.80\times$ Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that LISTEN preserves speech naturalness and conversational quality, while SPEAK retains conversational quality with some degradation in speech naturalness.