Real-Time Joint Audio-Video Generation by Parallel Adapter Composition

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the capability degradation caused by sequentially training streaming and few-step sampling for real-time audio-visual generation. To overcome this, we propose parallel training of causal and few-step adapters on a frozen diffusion Transformer backbone. Exploiting the approximate orthogonality of their update directions, the adapters are combined via direct addition based on model merging principles, thereby avoiding the interference inherent in chained fine-tuning without requiring explicit constraints. By integrating block-autoregressive attention with LoRA, the proposed method achieves continuous real-time generation at approximately 26 fps at 480×832 resolution. The resulting image quality is comparable to that of bidirectional teacher models and significantly surpasses conventional baselines.
📝 Abstract
Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap. Conventionally, the streaming video literature obtains both capabilities from a chained pipeline. It first distills a bidirectional teacher into a causal student, then into a few-step one, or proceeds in reverse order. Each stage of such a chain fine-tunes the weights the previous one produced, so a later objective can undo an earlier capability. Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel. A causal adapter is trained against the frozen backbone, and an off-the-shelf few-step adapter provides the few-step capability. As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training. Orthogonal updates should combine without interfering, so parallel composition is a direct sum. The two adapters are simply added at inference, with no joint training, yielding few-step, streaming audio-video whose image quality tracks the bidirectional teacher. Compared to the chained baselines, the composed model matches or beats them on most metrics, making parallel composition a practical approach. The resulting streaming system generates joint audio-video in real time, $\approx$26 fps at $480\times832$ without quantization, and sustains 30 s of continuous generation with stable image quality.
Problem

Research questions and friction points this paper is trying to address.

joint audio-video generation
real-time streaming
few-step sampling
causal attention
chained fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Adapter Composition
Audio-Video Diffusion Transformer
Model Merging
Block-Autoregressive Attention
Few-Step Sampling