🤖 AI Summary
This study addresses the tendency of existing video generation methods to produce trajectory drift and lose visual diversity during camera path control. To overcome these limitations, this work proposes a training-free particle filtering framework that steers pretrained models through sequential Monte Carlo (SMC) guidance combined with diffusion-based reward mechanisms. The core innovation lies in introducing a particle resampling-based local refinement stage, which effectively resolves the exploration–convergence trade-off inherent in sampling guidance and enables synergistic global and local optimization. Without requiring retraining, the proposed method precisely adheres to specified camera trajectories while significantly improving both trajectory fidelity and visual quality. Furthermore, it effectively suppresses drift artifacts, offering a robust and practical solution for controllable video generation.
📝 Abstract
We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid the need to construct large camera-annotated datasets by steering pretrained models toward the desired camera motion at test time. This enables the generation of camera-controlled video data that can subsequently be used to train camera-conditioned video diffusion models. Existing sampling-based guidance approaches often suffer from unstable trajectories: they either explore too broadly and fail to respect the target camera motion or collapse early and lose visual diversity over time. We introduce a global-local refinement framework for diffusion reward guidance, enabling accurate and consistent camera control during video generation. Our method builds on Sequential Monte-Carlo (SMC) guidance, but introduces a local refinement stage based on particle filtered resampling. Experiments show large improvements in camera trajectory adherence, reduced drift, and better visual quality, without requiring model retraining.