DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of imprecise action following and biased interaction prediction in robotic world models by constructing a multi-view, cross-embodied world model. Methodologically, it introduces image-space action-conditioned rendering and offline geometric calibration to enhance action alignment. Additionally, a counterfactual post-training mechanism based on human feedback is designed to broaden interaction coverage, while a video reward model is incorporated to optimize the physical plausibility of predictions. Experimental results demonstrate that the proposed approach achieves state-of-the-art action-following accuracy on the AgiBot dataset, reducing the interaction defect rate to 6.25%. Furthermore, this work secured first place in the 2026 AgiBot World Challenge.
πŸ“ Abstract
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at https://brave-eai.github.io/DreamTrue.
Problem

Research questions and friction points this paper is trying to address.

robot world model
action following
video prediction
cross-embodiment
interaction coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Robot World Model
Counterfactual Post-Training
Cross-embodiment
Geometric Calibration
Video Reward Model
πŸ”Ž Similar Papers