$R^2$-WAM: Repair-and-Reject Post-Training for World Action Models

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the issue of policy optimization being misled by inconsistencies between visual predictions and input actions in world action models. We propose a two-stage "repair-reject" post-training framework that first rectifies predictive consistency via kinematic alignment scoring, then performs negative fine-tuning by filtering suboptimal actions based on imagined trajectories, enabling policy optimization without additional environmental interactions. This approach extends video prediction from representation learning to consequence-based policy evaluation, introducing a novel mechanism that effectively mitigates the adverse effects of predictive bias. Evaluated on the RoboTwin 2.0 benchmark, our method achieves an average success rate of 93.8%, and 87.5% on a real-world cloth-folding task, significantly outperforming existing baselines.
πŸ“ Abstract
World Action Models (WAMs) emerge as a promising foundation for policy refinement by predicting the consequences of sampled actions. However, visually plausible predictions can mislead policy refinement if they fail to reflect the input actions. To address this mismatch, we introduce $R^2$-WAM, a two-stage repair-and-reject post-training framework that first improves the consistency of predicted futures with input actions, then uses these futures to select inferior action samples for negative fine-tuning. The repair stage grounds imagination in observed robot behavior through a kinematic alignment score that measures agreement between predicted and demonstrated motion, enabling the predicted video to faithfully reflect its input actions. Using the repaired video model, the rejection stage compares imagined outcomes of sampled and demonstrated actions, selectively applying negative fine-tuning to samples whose predicted task progress falls below the demonstrated reference by a prescribed margin. Together, the two stages extend video prediction from representation learning to consequence-based policy refinement without additional environment interaction or changes to the inference procedure. $R^2$-WAM achieves 93.8% average success on RoboTwin 2.0 across clean and randomized settings. On the long-horizon real-world Fold Shirt task, it achieves 87.5% average success, compared with 0% for Fast-WAM.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
policy refinement
action-conditioned video prediction
prediction-action mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Repair-and-Reject Post-Training
Kinematic Alignment Score
Negative Fine-Tuning
Policy Refinement
πŸ”Ž Similar Papers
No similar papers found.