Rethinking Representations for World-Action Modeling

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of unified design standards in the representation spaces of existing world action models, where reliance on reconstruction or pretrained features inadequately supports policy learning. We propose ReWAM, a representation-centric world action model built upon DINO features. Its core innovation lies in introducing a temporal representation bottleneck coupled with a gradient routing mechanism: by backpropagating action loss gradients exclusively to the bottleneck layer, the policy actively defines the encoded content while the world model focuses on dynamics evolution, thereby achieving efficient decoupling of control and prediction. Without requiring generative video pretraining, this approach attains a 93.6% success rate on RoboTwin 2.0 and an average score of 12.29 on RoboDojo, utilizing only approximately 600 hours of embodied data.
📝 Abstract
World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Action-Grounded Representation Shaping routes only action-loss gradients to the bottleneck, thereby letting the policy shape what the representation encodes while the world model learns how it evolves. Without generative video pre-training, ReWAM achieves 93.6% success on RoboTwin 2.0. On RoboDojo, it achieves an average score of 12.29 and a success rate of 8.28% using approximately 600 hours of embodied pre-training data.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Representation Learning
Robot Policy
Dynamics Modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Model
Representation Learning
DINO Features
Temporal Representation Bottleneck
Action-Grounded Representation Shaping
🔎 Similar Papers
No similar papers found.