🤖 AI Summary
This study addresses the limitation of existing camera-controlled video generation models in handling dynamic scenes, which frequently results in anomalous object motion or visual degradation. To overcome this, we propose a lightweight dynamics injection method that exploits the asymmetry between global camera motion and local object dynamics. Specifically, learnable scene tokens are introduced to capture local dynamics from sparse trajectories via cross-attention mechanisms. These tokens are then integrated into a frozen base model through an adaptive interface, enabling test-time adaptation without full fine-tuning. Extensive evaluations on the VBench2 and WorldScore benchmarks demonstrate that the proposed approach significantly outperforms both LoRA-based and full fine-tuning baselines, achieving a superior balance between dynamic content generation and precise camera control.
📝 Abstract
Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/