CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of generation quality in video diffusion model alignment within latent spaces, where fixed reward models induce distribution shift and implicit reward hacking. We propose CoRe, a framework that treats the generator and reward model as a dynamically interacting system. By continuously refitting the reward model while anchoring it to genuine human preferences, CoRe replaces static proxy optimization with a co-evolutionary mechanism, preventing the generator from exploiting spurious high rewards by deviating from the data distribution. This work is the first to identify distribution shift as the core cause of latent-space reward hacking. Experiments on Wan2.1-T2V-1.3B demonstrate that our approach effectively prevents quality collapse, achieving generation performance significantly superior to both the pretrained baseline and existing alignment methods.
📝 Abstract
Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.
Problem

Research questions and friction points this paper is trying to address.

latent reward hacking
video diffusion models
distributional escape
latent reward models
alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Reward Hacking
Distributional Escape
Co-Evolving Reward Models
Video Diffusion Models
Latent Space Alignment
🔎 Similar Papers
No similar papers found.