RobotAPO: Adversarial Physics Preference Optimization for Robotic Manipulation Video Generation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of downstream execution failures caused by physical violations in robot manipulation video generation by proposing the RobotAPO framework. Methodologically, it constructs the AgiBot-PhysPref dataset and introduces an adversarial counterfactual proposer to explore physical failure boundaries. By integrating continuous flow matching denoising spaces with adversarial preference optimization, the framework achieves efficient training under a pure prompt-based inference interface. Experimental results demonstrate that this approach significantly enhances the physical consistency of generated videos, yielding a 37.4% relative improvement in real-world robotic task success rates. Ultimately, this work establishes a new paradigm for physics-grounded, video-driven policy learning.
📝 Abstract
Robotic manipulation videos are increasingly used as visual plans for embodied agents, but optimizing purely for visual plausibility often fails to capture the fragile physical manifold of real-world interactions. Even minor physics-violating errors at the interaction boundary, such as interpenetration or premature object motion, can completely invalidate the inferred timing and pose needed for downstream execution. Because standard supervised fine-tuning lacks the direct pressure to penalize these localized failures, we introduce AgiBot-PhysPref. This rigorously curated 10,000-sample preference dataset isolates condition-matched physics violations, turning the generator's own failure distribution into a foundational signal for physical consistency. Building upon this, we propose RobotAPO, an adversarial physics preference optimization framework operating in the continuous flow-matching denoising space. To prevent the policy from merely memorizing static curated failures, RobotAPO employs a lightweight adversarial counterfactual proposer that learns a condition-dependent, physical-failure-biased direction in denoising space. This encourages the model to explore and better respect the physical interaction boundary, all while maintaining a pure prompt-and-reference inference interface without requiring external structural conditioning. Comprehensive evaluations demonstrate that explicitly correcting these localized physics violations improves downstream robot execution from generated videos. On held-out AgiBot conditions, RobotAPO outperforms the strongest controlled internal baseline in physical consistency by 6.8% hard score and 10.0% soft score. Crucially, in real-robot replay, it translates these physical-consistency gains into a 37.4% relative improvement in task success over the strongest controlled internal baseline.
Problem

Research questions and friction points this paper is trying to address.

robotic manipulation video generation
physical consistency
physics violations
embodied agents
downstream execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Physics Preference Optimization
Flow-matching Denoising Space
Robotic Manipulation Video Generation
Preference Dataset
Physical Consistency
🔎 Similar Papers
No similar papers found.
K
Kerui Li
Institute of Automation, Chinese Academy of Sciences
Z
Zhe Jing
Beijing Institute of Technology
C
Chenyi Huang
Institute of Automation, Chinese Academy of Sciences
X
Xiaofeng Wang
GigaAI
Z
Zheng Zhu
GigaAI
H
Haoming Cui
Institute of Automation, Chinese Academy of Sciences
Huaibo Huang
Huaibo Huang
NLPR, MAIS, CASIA
Computer VisionGenerative ModelsLow-level VisionFace Recognition