🤖 AI Summary
This study addresses the limitation of visual world model planning, where directly scoring actions against final goals often leads to local optima and fails in tasks requiring initial movement away from the target. To overcome this, we propose an anchored planning strategy that, under a frozen world model, retrieves short-horizon intermediate observations via experience replay to replace the single final goal, dynamically adjusting the target horizon for action synthesis and ranking. Our analysis further reveals that lower prediction error does not necessarily yield superior control. This approach significantly enhances long-horizon control capabilities without additional training. It substantially outperforms the original LeWM planner on benchmarks such as Cube and PushT, successfully accomplishing previously intractable tasks solely by modifying the planning objective.
📝 Abstract
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed experience both produce these gains. We introduce Anchored Planning, which retrieves a recorded segment whose start and end resemble the current and goal observations, then aims at an observation shortly after its start. The frozen model scores actions toward this target from the current state. Without additional training, planning toward observed targets outperforms the released LeWM planner on every task in our long-range evaluation. Additional final-goal search falls short of the same gains. Lower successor-prediction error need not translate into better control. Success also depends on how far ahead the target is placed and on shrinking the retrieval span as execution advances. Changing only the target lets the same frozen model and planner reach goals that final-goal scoring misses.