Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of fixed reward routing and weighting schemes in joint audio-video diffusion models to adapt to training dynamics by proposing an adaptive reward routing framework. Methodologically, it dynamically adjusts update positions and coordinates multi-objective rewards within a forward reinforcement learning paradigm. The framework innovatively introduces cross-modal influence-guided routing alongside a preference-preserving, modality-aware reweighting mechanism. Furthermore, training is optimized by integrating bidirectional cross-attention responses, token-level loss reweighting, and gradient scaling techniques. Experimental results demonstrate that the proposed approach significantly outperforms existing strong baselines across multiple dimensions, including unimodal generation quality, semantic consistency, and audio-visual synchronization.
📝 Abstract
Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.
Problem

Research questions and friction points this paper is trying to address.

multi-reward optimization
audio-video diffusion models
reinforcement learning
reward routing
reward coordination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Reward Routing
Forward-Process RL
Cross-Modal Influence-Guided Routing
Preference-Preserving Reweighting
Joint Audio-Video Diffusion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Songlin Yang
MMLab@HKUST, The Hong Kong University of Science and Technology
X
Xiaotong Zhao
Tencent Video
J
Jiacheng Zhang
The University of Hong Kong
Zhe Wang
Zhe Wang
The Hong Kong University of Science and Technology
Atmospheric chemistryHeterogeneous ChemistrySOA formationCloud-aerosol-gas interaction
T
Toyota Li
Tencent Video
Eric Liu
Eric Liu
University of Toronto
SecurityCompilersFuzzing
A
Alan Zhao
Tencent Video
Anyi Rao
Anyi Rao
Assistant Professor, HKUST
Human AIAI for CreativityGenerative AIContent CreationFilm