🤖 AI Summary
This study addresses the limitation of full-image reconstruction objectives in world models, where irrelevant visual content tends to dominate representations and interfere with dynamics learning. To mitigate this, we propose a dynamics-aware adaptive reconstruction mechanism that employs self-supervised attention routing to guide a dual-stream decoder toward reconstructing task-relevant regions. Furthermore, we introduce a stop-gradient barrier to decouple visual supervision from state prediction, preventing non-predictive information from corrupting the latent space. Evaluated on the DeepMind Control benchmark under both random frame and sequential video distractions, our approach achieves state-of-the-art performance by faithfully preserving essential state attributes while effectively suppressing noise.
📝 Abstract
A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant information. We propose StarWM, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies. A dual-stream decoder then restricts reconstruction to the attended regions, with stop-gradient barriers preventing interference between the two objectives. These components allows reconstruction to supervise the visual content of attended regions without contaminating the latent with non-predictive information. On DeepMind Control with dynamic video backgrounds, default (reward-free) StarWM achieves the strongest performance under random-frame distractors and substantially outperforms reconstruction-based baselines under sequential video. In addition, its reward-augmented variant matches or exceeds reconstruction-free methods on sequential video, achieving the highest overall return across all distractor regimes. Mechanistic probing confirms StarWM preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.