π€ AI Summary
This study addresses the challenges of imprecise action following and biased interaction prediction in robotic world models by constructing a multi-view, cross-embodied world model. Methodologically, it introduces image-space action-conditioned rendering and offline geometric calibration to enhance action alignment. Additionally, a counterfactual post-training mechanism based on human feedback is designed to broaden interaction coverage, while a video reward model is incorporated to optimize the physical plausibility of predictions. Experimental results demonstrate that the proposed approach achieves state-of-the-art action-following accuracy on the AgiBot dataset, reducing the interaction defect rate to 6.25%. Furthermore, this work secured first place in the 2026 AgiBot World Challenge.
π Abstract
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at https://brave-eai.github.io/DreamTrue.