DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing camera-controlled video generation models in handling dynamic scenes, which frequently results in anomalous object motion or visual degradation. To overcome this, we propose a lightweight dynamics injection method that exploits the asymmetry between global camera motion and local object dynamics. Specifically, learnable scene tokens are introduced to capture local dynamics from sparse trajectories via cross-attention mechanisms. These tokens are then integrated into a frozen base model through an adaptive interface, enabling test-time adaptation without full fine-tuning. Extensive evaluations on the VBench2 and WorldScore benchmarks demonstrate that the proposed approach significantly outperforms both LoRA-based and full fine-tuning baselines, achieving a superior balance between dynamic content generation and precise camera control.
📝 Abstract
Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation. Project website: https://glab-caltech.github.io/dynatokens/
Problem

Research questions and friction points this paper is trying to address.

video generation
camera control
scene dynamics
dynamic scenes
motion modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

DynaTokens
Test-time Adaptation
Camera-controlled Video Generation
Scene Dynamics
Cross-attention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziqi Ma
California Institute of Technology
H
Hongqiao Chen
California Institute of Technology
Georgia Gkioxari
Georgia Gkioxari
Caltech, Meta AI
Computer VisionMachine LearningArtificial Intelligence