Achieve What You Imagined: Learning to Align Actions with Visual Plans

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistency between visual predictions and generated action outcomes in world action models by proposing an unsupervised optimization framework based on cross-model prediction discrepancies. The method treats visual predictions as goal proposals, constructs feedback signals using a frozen action-conditioned world model, and achieves action alignment through flow policy optimization. A core innovation lies in eliminating the need for online interaction or additional reward training, relying solely on cross-model prediction discrepancies as a self-supervised optimization signal. Evaluated across four real-world UR5 robotic manipulation tasks, the proposed approach significantly improves success rates from 43.4% to 75.1%, substantially outperforming existing baseline methods.
📝 Abstract
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences that reliably achieve them. Consequently, we treat the WAM-generated visual prediction as a goal-conditioned visual proposal rather than a directly executable plan. We use a frozen action-conditioned world model to predict action-conditioned consequences and construct feedback based on consistency between the two future predictions and alignment with the terminal goal. Leveraging this feedback, we employ Flow Policy Optimization (FPO) to optimize the action head of the WAM. This framework avoids online robot interaction and additional training of task-specific reward models. Across four real-world UR5 manipulation tasks, our method increases the mean success rate from 43.4% to 75.1%, compared with 61.4% for $\pi_{0.5}$. These results show that cross-model prediction discrepancy can provide useful feedback for improving robot policies under the evaluated manipulation tasks. Website: https://imagine-to-achieve.github.io/
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Visual-Action Alignment
Prediction Discrepancy
Robot Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Models
Flow Policy Optimization
Visual Planning
Goal-Conditioned Policy
Cross-Model Consistency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuheng Qiao
KTH
Z
Ziran Wei
KTH
X
Xiaohan Wang
Beihang University
D
Daqiang Guo
The Hong Kong University of Science and Technology (Guangzhou)
Yichen Luo
Yichen Luo
PhD Candidate, Department of Computer Science, University College London
Multi-Agent SystemsBlockchainGreen Finance
Zhibo Pang
Zhibo Pang
ABB Corporate Research, and KTH Royal Institute of Technology, Sweden
RoboticsAICloudWirelessIndustrial Automation
P
Peng Zhou
Great Bay University
S
Sichao Liu
KTH