Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the context mismatch between causal modeling and distillation in few-step streaming audio-visual generation by proposing a two-stage post-training framework. First, it introduces Causal Self-Flow based on historical noise mixing to achieve self-supervised semantic alignment. Second, it performs context-aligned autoregressive Distribution Matching Distillation (DMD) under a unified causal mask, combined with a block-conditional KL objective, enabling few-step distillation without additional consistency losses. Experimental results demonstrate that this approach improves visual and motion quality at 480p resolution by 57% and 45%, respectively, over baselines, while its high-resolution generation performance surpasses that of bidirectional models.
📝 Abstract
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step $1664\times960$ generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
Problem

Research questions and friction points this paper is trying to address.

streaming audio-video generation
causal modeling
step distillation
context mismatch
teacher forcing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Multimodal Generation
Causal Self-Flow
Distribution Matching Distillation
Few-Step Distillation
Post-Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xingtong Ge
Xingtong Ge
Hong Kong University of Science and Technology, SenseTime, Beijing Institute of Technology
Diffusion modelsImage/Video CompressionGaussian Splatting
Yutong Wang
Yutong Wang
The University of Sydney
Computer VisionCrowd Counting
L
Lunjie Zhu
The Hong Kong University of Science and Technology
Haitao Lin
Haitao Lin
Westlake University & Zhejiang University
Generative ModelsGeometric Deep LearningAI4Science
F
Fangyu Lin
The Hong Kong University of Science and Technology
Yushi Huang
Yushi Huang
Hong Kong University of Science and Technology
Efficient AI
X
Xin Zhang
Vivix Group Limited
Y
Yi Zhang
Vivix Group Limited
Y
Yu Liu
Vivix Group Limited
J
Jun Zhang
The Hong Kong University of Science and Technology