🤖 AI Summary
This study addresses the inefficiency of existing robotic planning methods that fully restart upon execution feedback invalidating visual predictions, despite the underlying task structure remaining valid. To overcome this, we propose a revisable temporal planning framework introducing the first visual plan revision strategy based on intermediate state preservation and a learned bridging mechanism. By integrating world action models, generative breakpoint resumption, adaptive policy selection, and temporally aware history, our approach recovers from intermediate states and adapts to new observations for efficient closed-loop control. This method eliminates the computational overhead of complete replanning. Extensive evaluations demonstrate its effectiveness, achieving average task success rates of 48.6% and 84.8% on the RoboMME and RMBench benchmarks, respectively.
📝 Abstract
World-action models use predicted visual futures to condition robot actions, yet execution feedback can invalidate parts of a prediction while leaving its task structure useful. We propose Revisable Temporal Planning (RTP), which maintains the visual future as a persistent action condition and revises it after feedback. Its central mechanism is a learned revision bridge: it resumes an intermediate state saved during visual generation and adapts its continuation to current observations. Visual and action supervision connect this revision to subsequent control. Time-aware history supplies observed evidence, and an adaptive policy selects retention, bridge revision, or fresh replanning from new noise before decoding the next action. On RoboMME and RMBench, RTP achieves task-averaged success rates of 48.6% and 84.8%, respectively. Matched comparisons support learned continuation; estimated checkpoint-source and action-prefix effects are positive but less precisely resolved. These results connect feedback-driven visual-plan revision to closed-loop task performance. Project Page: https://PLACEHOLDER.github.io/RTP/