LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses 3D inconsistency issues in long-horizon camera-controlled video generation, such as permanent object loss and scene structure drift. To tackle these challenges, we propose LoGo, a post-training reinforcement learning framework that integrates global and local reward signals to enable fine-grained credit assignment, significantly enhancing 3D consistency while preserving camera-following capability and video quality. Additionally, we introduce TrajectoryBench, a benchmark designed for systematically evaluating generative performance under complex, extended camera trajectories. Experimental results demonstrate that LoGo substantially outperforms existing methods on both DL3DV and TrajectoryBench, effectively suppressing local object displacement, visual artifacts, and global scene drift. Furthermore, the proposed approach exhibits strong generalization across diverse foundation models.
📝 Abstract
Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Existing post-training techniques, which assign a single scalar reward to the entire generation, are poorly suited to correcting these inconsistencies over long horizons. We introduce LoGo, which blends global and spatially localized rewards for camera-controlled video models. The local reward provides fine-grained credit assignment, which substantially improves 3D consistency, while the global reward preserves camera following and video quality. Across three base models, LoGo shows a clear advantage on DL3DV and TrajectoryBench, a new benchmark for long-horizon, complex-camera-control generation that current evaluations lack. LoGo effectively reduces local object shifts, artifacts, and global scene changes, illustrating the importance of credit assignment in post-training video models. Project website: https://ziqi-ma.github.io/logo-website/
Problem

Research questions and friction points this paper is trying to address.

long-horizon video generation
3D inconsistency
camera-controlled video models
reward assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local-Global Rewards
Credit Assignment
3D Consistency
Long-Horizon Video Generation
Camera-Controlled Video Models
🔎 Similar Papers
No similar papers found.