WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high costs of robot learning, imprecise action following in video models, and the difficulty of sharing data across embodiments by proposing a visual simulator that decouples dynamics learning from action grounding. Methodologically, it introduces an image-space action representation to unify control interfaces, combining relational regularization with few-step distillation for efficient causal reasoning. The model learns dynamics from 10,000 hours of unlabeled videos and achieves grounding using 1,000 hours of trajectory data, further incorporating multi-view training and failure augmentation. Experimental results demonstrate that the Intersection over Union (IoU) for failure trajectories improves by 0.16, prediction accuracy reaches 74%, and zero-shot transfer increases task success rates by up to 21.4%.
📝 Abstract
Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{https://zhengsh123.github.io/WorldLine/}{project page}.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation
visual simulation
video generation
action grounding
cross-embodiment transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action-driven visual simulation
Decoupled dynamics learning
Image-space action representation
Few-step distillation
Robotic manipulation
🔎 Similar Papers
No similar papers found.