Institution profile

Zuoyebang Education Technology

Industry researchasia · cn
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Sep 24, 2026

This study addresses the lack of adaptive mechanisms in existing full-duplex speech dialogue models, which struggle to meet dynamic interaction demands. Building upon Qwen3-Omni, this work proposes a three-tier collaborative adaptive full-duplex architecture. Methodologically, it introduces a compact token protocol and training-free runtime control to support dynamic window prediction and non-blocking inference. A three-stage progressive curriculum learning strategy is designed to decouple behavioral decision-making. Furthermore, the framework integrates dual-stream alignment, logits bias, GRPO reinforcement learning, and in-flight external reasoning techniques. Experimental results demonstrate that the proposed method surpasses state-of-the-art performance across multiple metrics on Full-Duplex-Bench v1 and v1.5, achieving a leading score of 72.9 on HumDial-FDBench.

0 citationsRead paper

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

Jun 30, 2026

Flow Matching in speech synthesis suffers from high inference latency and timbre leakage. This work proposes a unified guidance framework that, for the first time, jointly integrates data-level and model-level guidance. By leveraging heterogeneous data augmentation to disentangle linguistic content from acoustic residuals, and combining trajectory correction with an intrinsic guidance objective, the method distills conditional information directly into network weights to optimize the inference trajectory. Notably, it eliminates the need for Classifier-Free Guidance, substantially reducing computational overhead. The approach achieves nearly threefold faster inference while preserving high timbre fidelity and significantly outperforms state-of-the-art baselines in speaker similarity.

0 citationsRead paper
Recent publications

Latest Papers

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Sep 24, 2026

This study addresses the lack of adaptive mechanisms in existing full-duplex speech dialogue models, which struggle to meet dynamic interaction demands. Building upon Qwen3-Omni, this work proposes a three-tier collaborative adaptive full-duplex architecture. Methodologically, it introduces a compact token protocol and training-free runtime control to support dynamic window prediction and non-blocking inference. A three-stage progressive curriculum learning strategy is designed to decouple behavioral decision-making. Furthermore, the framework integrates dual-stream alignment, logits bias, GRPO reinforcement learning, and in-flight external reasoning techniques. Experimental results demonstrate that the proposed method surpasses state-of-the-art performance across multiple metrics on Full-Duplex-Bench v1 and v1.5, achieving a leading score of 72.9 on HumDial-FDBench.

0 citationsRead paper

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

Jun 30, 2026

Flow Matching in speech synthesis suffers from high inference latency and timbre leakage. This work proposes a unified guidance framework that, for the first time, jointly integrates data-level and model-level guidance. By leveraging heterogeneous data augmentation to disentangle linguistic content from acoustic residuals, and combining trajectory correction with an intrinsic guidance objective, the method distills conditional information directly into network weights to optimize the inference trajectory. Notably, it eliminates the need for Classifier-Free Guidance, substantially reducing computational overhead. The approach achieves nearly threefold faster inference while preserving high timbre fidelity and significantly outperforms state-of-the-art baselines in speaker similarity.

0 citationsRead paper