ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

πŸ“… 2026-04-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing vision-language value models struggle to capture temporal dynamics, leading to unreliable value estimation in long-horizon robotic tasks. This work proposes ViVa, the first approach to incorporate spatiotemporal priors from pretrained video generation models into reinforcement learning. By fusing visual, linguistic, and proprioceptive inputs, ViVa jointly predicts future states and current values, enabling value modeling grounded in dynamic prediction rather than static snapshots. This formulation intrinsically couples value estimation with the agent’s ability to anticipate the consequences of its actions. Evaluated on real-world tasks such as box assembly, ViVa yields more accurate and reliable value signals and demonstrates strong generalization to novel objects.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionComputer Vision: Language and VisionIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsResponsible Web: Machine-in-the-loop, human agency and autonomy
πŸ“ Abstract
Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback. Reinforcement learning addresses this via value functions, which assess task progress and guide policy improvement. However, existing value models built on vision-language models (VLMs) struggle to capture temporal dynamics, undermining reliable value estimation in long-horizon tasks. In this paper, we propose ViVa, a video-generative value model that repurposes a pretrained video generator for value estimation. Taking the current observation and robot proprioception as input, ViVa jointly predicts future proprioception and a scalar value for the current state. By leveraging the spatiotemporal priors of a pretrained video generator, our approach grounds value estimation in anticipated embodiment dynamics, moving beyond static snapshots to intrinsically couple value with foresight. Integrated into RECAP, ViVa delivers substantial improvements on real-world box assembly. Qualitative analysis across all three tasks confirms that ViVa produces more reliable value signals, accurately reflecting task progress. By leveraging spatiotemporal priors from video corpora, ViVa also generalizes to novel objects, highlighting the promise of video-generative models for value estimation.
Problem

Research questions and friction points this paper is trying to address.

value estimation
temporal dynamics
robot reinforcement learning
vision-language models
long-horizon tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

video-generative value model
spatiotemporal priors
reinforcement learning
embodiment dynamics
vision-language-action
πŸ”Ž Similar Papers