🤖 AI Summary
This study addresses the problem of downstream execution failures caused by physical violations in robot manipulation video generation by proposing the RobotAPO framework. Methodologically, it constructs the AgiBot-PhysPref dataset and introduces an adversarial counterfactual proposer to explore physical failure boundaries. By integrating continuous flow matching denoising spaces with adversarial preference optimization, the framework achieves efficient training under a pure prompt-based inference interface. Experimental results demonstrate that this approach significantly enhances the physical consistency of generated videos, yielding a 37.4% relative improvement in real-world robotic task success rates. Ultimately, this work establishes a new paradigm for physics-grounded, video-driven policy learning.
📝 Abstract
Robotic manipulation videos are increasingly used as visual plans for embodied agents, but optimizing purely for visual plausibility often fails to capture the fragile physical manifold of real-world interactions. Even minor physics-violating errors at the interaction boundary, such as interpenetration or premature object motion, can completely invalidate the inferred timing and pose needed for downstream execution. Because standard supervised fine-tuning lacks the direct pressure to penalize these localized failures, we introduce AgiBot-PhysPref. This rigorously curated 10,000-sample preference dataset isolates condition-matched physics violations, turning the generator's own failure distribution into a foundational signal for physical consistency. Building upon this, we propose RobotAPO, an adversarial physics preference optimization framework operating in the continuous flow-matching denoising space. To prevent the policy from merely memorizing static curated failures, RobotAPO employs a lightweight adversarial counterfactual proposer that learns a condition-dependent, physical-failure-biased direction in denoising space. This encourages the model to explore and better respect the physical interaction boundary, all while maintaining a pure prompt-and-reference inference interface without requiring external structural conditioning. Comprehensive evaluations demonstrate that explicitly correcting these localized physics violations improves downstream robot execution from generated videos. On held-out AgiBot conditions, RobotAPO outperforms the strongest controlled internal baseline in physical consistency by 6.8% hard score and 10.0% soft score. Crucially, in real-robot replay, it translates these physical-consistency gains into a 37.4% relative improvement in task success over the strongest controlled internal baseline.