🤖 AI Summary
This study addresses the inability of fixed reward routing and weighting schemes in joint audio-video diffusion models to adapt to training dynamics by proposing an adaptive reward routing framework. Methodologically, it dynamically adjusts update positions and coordinates multi-objective rewards within a forward reinforcement learning paradigm. The framework innovatively introduces cross-modal influence-guided routing alongside a preference-preserving, modality-aware reweighting mechanism. Furthermore, training is optimized by integrating bidirectional cross-attention responses, token-level loss reweighting, and gradient scaling techniques. Experimental results demonstrate that the proposed approach significantly outperforms existing strong baselines across multiple dimensions, including unimodal generation quality, semantic consistency, and audio-visual synchronization.
📝 Abstract
Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.