Token-Level Video Reinforcement Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of holistic scalar rewards in reinforcement learning for video generation, which fail to localize errors and consequently perturb satisfactory tokens while neglecting those requiring correction. To overcome this, we propose TVRL, a framework that introduces the first token-level credit assignment mechanism based on input gradient magnitudes from frozen vision-language models. By integrating GRPO with an SDE sampler to reweight dense denoising transition log-probabilities, TVRL achieves fine-grained mapping from video-level rewards to token-level optimization. Experimental results demonstrate that our method attains a score of 57.69 on VBench-2.0, outperforming the baseline by 3.60 points, and consistently surpasses standard GRPO across diverse sampler and reward model configurations.
📝 Abstract
Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.
Problem

Research questions and friction points this paper is trying to address.

Video Generation
Reinforcement Learning
Token-Level Reward
Credit Assignment
Scalar Reward Limitation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-Level Reinforcement Learning
Video Generation
Credit Assignment
Vision-Language Model Gradients
Group Relative Policy Optimization
🔎 Similar Papers
No similar papers found.