LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation
This study addresses 3D inconsistency issues in long-horizon camera-controlled video generation, such as permanent object loss and scene structure drift. To tackle these challenges, we propose LoGo, a post-training reinforcement learning framework that integrates global and local reward signals to enable fine-grained credit assignment, significantly enhancing 3D consistency while preserving camera-following capability and video quality. Additionally, we introduce TrajectoryBench, a benchmark designed for systematically evaluating generative performance under complex, extended camera trajectories. Experimental results demonstrate that LoGo substantially outperforms existing methods on both DL3DV and TrajectoryBench, effectively suppressing local object displacement, visual artifacts, and global scene drift. Furthermore, the proposed approach exhibits strong generalization across diverse foundation models.