Institution profile

CPII

Research institution
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

Oct 02, 2026

This study addresses the challenges of insufficient visual quality and semantic misalignment in streaming video generation by proposing a distillation framework based on unified joint-marginal distribution matching. Methodologically, an image teacher model is leveraged to provide frame-level supervision for optimizing few-step generation. Furthermore, a LatentBridge mechanism is introduced to resolve cross-domain latent representation mismatches, combined with latent variant sampling to enhance dynamic representational capacity. Experimental results demonstrate that the proposed approach significantly improves both visual fidelity and text alignment while effectively preserving motion coherence, achieving a human preference rate exceeding 80%.

0 citationsRead paper

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

Sep 29, 2026

This study addresses the limitations of full-duplex speech models in long-context and multi-party interactions by extending the Moshi architecture to long-duration, multi-party, and Chinese-English bilingual scenarios. Methodologically, it proposes a codec frame-level bilingual end-to-end full-duplex modeling framework and introduces the first large-scale dataset and evaluation benchmark supporting the joint modeling of complex interaction features, including duration, overlap, and interruption. The project releases 57,600 hours of synthetic data alongside a real-recorded benchmark, MultiTalkBench. Experimental results demonstrate that the trained model significantly outperforms open-source baselines such as Moshi and MiniCPM-o on MultiTalkBench, achieving high coherence and precise responsiveness in extended multi-party dialogues.

0 citationsRead paper
Recent publications

Latest Papers

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

Oct 02, 2026

This study addresses the challenges of insufficient visual quality and semantic misalignment in streaming video generation by proposing a distillation framework based on unified joint-marginal distribution matching. Methodologically, an image teacher model is leveraged to provide frame-level supervision for optimizing few-step generation. Furthermore, a LatentBridge mechanism is introduced to resolve cross-domain latent representation mismatches, combined with latent variant sampling to enhance dynamic representational capacity. Experimental results demonstrate that the proposed approach significantly improves both visual fidelity and text alignment while effectively preserving motion coherence, achieving a human preference rate exceeding 80%.

0 citationsRead paper

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

Sep 29, 2026

This study addresses the limitations of full-duplex speech models in long-context and multi-party interactions by extending the Moshi architecture to long-duration, multi-party, and Chinese-English bilingual scenarios. Methodologically, it proposes a codec frame-level bilingual end-to-end full-duplex modeling framework and introduces the first large-scale dataset and evaluation benchmark supporting the joint modeling of complex interaction features, including duration, overlap, and interruption. The project releases 57,600 hours of synthetic data alongside a real-recorded benchmark, MultiTalkBench. Experimental results demonstrate that the trained model significantly outperforms open-source baselines such as Moshi and MiniCPM-o on MultiTalkBench, achieving high coherence and precise responsiveness in extended multi-party dialogues.

0 citationsRead paper