ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video world models struggle with precise cross-scenario action control due to loose action representations or reliance on structured signals. This work proposes “Shadow Pairs”—video pairs sharing identical dynamics but differing in appearance—and leverages cross-shadow prediction to disentangle appearance from dynamics, thereby learning a unified action representation. Without requiring action annotations or fine-tuning, the method enables any demonstration video to be reused as a high-fidelity, transferable action asset in novel environments. Built upon a large-scale shadow video corpus and a block-causal world model, the approach substantially outperforms current baselines across diverse dynamical scenarios, achieving an average blind-test win rate of 86% in long-horizon action replay tasks.
📝 Abstract
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io
Problem

Research questions and friction points this paper is trying to address.

video world models
any-action control
dynamics representation
action transfer
frame-level control
Innovation

Methods, ideas, or system contributions that make the work stand out.

shadow pairs
cross-shadow prediction
unified dynamics representation
interactive video world models
action transfer
🔎 Similar Papers
2024-01-15IEEE Transactions on Information Forensics and SecurityCitations: 0
2024-09-10arXiv.orgCitations: 0