FutureWorlds: Learning Robotic World Models from Alternative Futures

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key challenges in robotic world models, including the difficulty of translating surrogate predictions into effective learning signals, low candidate diversity, and cumbersome trajectory history maintenance. To this end, we propose a unified optimization framework that integrates a multimodal discrete autoregressive architecture with diverse beam search to generate candidate future scenarios balancing confidence and diversity. We introduce a candidate-specific bounded memory mechanism to ensure historical consistency between generation and scoring, alongside the MemSPO algorithm, which converts video rewards into group-relative advantages for policy optimization. Experimental results demonstrate that our approach significantly reduces LPIPS scores on benchmarks such as RT-1, substantially improving generation quality within only 200 update steps while achieving more precise motion prediction and consistent object states.
📝 Abstract
Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: https://github.com/Alexander-wu/FutureWorlds.
Problem

Research questions and friction points this paper is trying to address.

Robotic World Models
Alternative Futures
Reinforcement Learning
Candidate Construction
History Maintenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Robotic World Models
Diverse Beam Search
Bounded Memory
MemSPO
Reinforcement Learning