FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing world models that fail to distinguish visual dynamics between egocentric and wrist-mounted views, resulting in mismatched future supervision targets. We propose a dual-path independent prediction mechanism that designs view-specific supervision objectives, integrating scene evolution with local interaction states for precise control. Furthermore, observation access is decoupled from future prediction, ensuring auxiliary modules incur zero inference overhead by being employed exclusively during training. Built upon ActionDiT, our method jointly models tasks and interaction states through multimodal prediction using RGB, masks, and skeletons. Evaluations on the RoboTwin50 and LIBERO benchmarks demonstrate success rates of 94.2% and 99.2%, respectively, indicating significant improvements in fine-grained manipulation performance.
📝 Abstract
World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution, whereas wrist cameras move with the end effector, mixing local interaction changes with viewpoint shifts and self-occlusion. These contrasting predictive demands suggest that the two views may benefit from different future objectives. We introduce FutureDuet, which retains both views for control, while allowing each visual stream to receive a different future objective. For the main view, future RGB models task evolution, while interaction masks and robot skeletons focus supervision on task objects and robot motion. For the wrist stream, future latent prediction models short-horizon interaction changes without requiring pixel-level reconstruction. ActionDiT jointly reads the resulting Task State and Interaction State, combining scene-level progress with close-range interaction evidence. All auxiliary prediction modules are training-only, adding no inference overhead. FutureDuet achieves 94.2% clean and 94.1% randomized success on RoboTwin50 and 99.2% average success on LIBERO. The improvements are most pronounced on six RoboTwin50 tasks that require precise interaction, averaging gains of 9.2% and 12.8% over Fast-WAM in clean and randomized settings. Controlled studies further show complementary gains from separating the wrist pathway and designing future supervision separately for the two views.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
robot manipulation
multi-view observation
future supervision
visual dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Decoupled Future Supervision
ActionDiT
Latent Prediction
Multi-view Observations