Future Anchored Verification and Online Recovery for World Action Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of execution drift in world action models, which leads to task failure and remains difficult to recover from using existing monitoring approaches. To this end, we propose FAVOR, a framework whose core innovation lies in transforming predicted future frames into verifiable anchors. By integrating vision-language models (VLMs), FAVOR establishes a lightweight, closed-loop error correction mechanism that requires no policy modification. Specifically, an anchor verifier compares observed and expected states to detect deviations, while a VLM generates concise corrective instructions to guide the model back to its intended trajectory. Experiments demonstrate that FAVOR improves success rates to 98.10% and 72.98% on the LIBERO and LIBERO-Plus benchmarks, respectively, significantly enhancing the online recovery capability and robustness of world action models.
📝 Abstract
World action models (WAMs) have emerged as a promising paradigm for robotic manipulation. They act by first predicting how a task should be performed and then decoding the actions from that future. However, the remaining actions are invalid once execution drifts from the prediction. Simply replanning from the already out of distribution state rarely restores what the task still requires; existing execution monitors decide when to stop, but not what to restore. We observe that the answer is already in hand: the future the WAM predicted before acting depicts exactly the states it intended to pass through. We introduce FAVOR (Future Anchored Verification and Online Recovery), a lightweight framework that keeps these predicted frames as anchors and uses them for verification and recovery. An Anchor Verifier compares each observation with its anchor, together with the executed actions, to flag deviations that break the task. Anchor-Guided Recovery uses a vision-language model to turn the flagged anchor into a short corrective instruction. Under strengthened instruction guidance, the WAM executes this instruction to return to the intended future. It then resumes the task. FAVOR raises the task success of the base WAM from 97.85% to 98.10% on LIBERO and from 72.60% to 72.98% on LIBERO-Plus without modifying the policy.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Execution Drift
Online Recovery
Robotic Manipulation
Execution Monitoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Future Anchored Verification
Online Recovery
Vision-Language Model
Robotic Manipulation
🔎 Similar Papers
No similar papers found.