Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of controllable world model training on expensive, synchronously annotated action data. To overcome this limitation, it proposes extracting ego-motion bases from unannotated videos via PCA, enabling the construction of a composable action space and the generation of grounded control signals without additional training. The method further employs a frozen decoder tracking architecture with video DiT, achieving efficient training through online latent critic distillation from a teacher model. As key contributions, this work introduces a reference-free evaluation metric suite and demonstrates linear scaling alongside precise control over reverse and stationary actions. It maintains high generation quality while overcoming the inherent limitations of conventional benchmarks in handling inverse motions.
📝 Abstract
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure induced by egomotion to obtain grounded control signals directly. Using a method as simple as principal components analysis perform this, we find that the leading components provide signed, scalable, and composable throttle--yaw controls, although the method can recover only motion axes represented in the data. To prevent a high-capacity video DiT from exploiting pixel-level supervision, an online latent critic distils a frozen decoder--tracker--PCA (Principal Components Analysis) teacher without backpropagating through the decoder or tracker. Finally we critique the use of video generation metrics to evaluate WMs and introduce an example of an alternative, reference-free evaluation method. We measure \textit{controllability}, \textit{plausibility}, \textit{conjuring} (creating objects out of thin air) and \textit{geometric integrity}, revealing failures that conventional video metrics miss. We show that most baselines follow familiar action directions but struggle to reverse or remain stationary. Our model handles both while retaining compositional control and generation quality. Despite backwards actions being less than $1\%$ of our training data, we find that the model learns to reverse, scale its response linearly, and compose throttle with steering, all simply by learning through a grounded action space.
Problem

Research questions and friction points this paper is trying to address.

World Models
Action Annotations
Unsupervised Video
Egomotion
Controllability Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Egomotion Basis
Principal Components Analysis
Video Diffusion Transformer
Controllability Evaluation
🔎 Similar Papers