Copper-Policy: Focus on the Representation for Robust Robot Manipulation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational costs of pixel-level world action models by proposing Copper-Policy, a framework that abandons predefined goal spaces and pixel reconstruction paradigms. Instead, it learns compact world representations through temporally joint embedding prediction and jointly optimizes representation learning with policy decoding. Coupled with an efficient GPU-parallelized architecture, this approach significantly reduces training overhead. Experimental results demonstrate that Copper-Policy accelerates training by sixfold while outperforming existing methods on both the RoboTwin and LIBERO-Plus benchmarks. Furthermore, the proposed framework exhibits superior robustness and generalization capabilities in real-world robotic tasks.
📝 Abstract
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce Copper-Policy, which learns a compact World representation with the policy rather than relying on a predefined target space. Through temporal joint-embedding prediction, it predicts future observation embeddings conditioned on task intention without reconstructing pixels. This prediction and action decoding shape the representation jointly, while the policy retains access to current-frame spatial detail for execution. Representation analyses show that the learned features better separate task-driven change from perturbations and provide complementary information for control. Compact prediction targets reduce training tokens per sample, enabling a 2B-parameter model trained in 9.67 hours on 8$\times$ RTX 5090 GPUs and 6$\times$ faster than Fast-WAM on matched A100 GPUs. Copper-Policy outperforms every compared method without embodied pretraining on RoboTwin and several embodied-pretrained VLAs on LIBERO-Plus (80.85%). On three challenging real-robot tasks, it performs comparably to $\pi_{0.5}$ and attains a higher average score. Together, these results show that Copper-Policy combines strong control performance with efficient training.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
robot manipulation
representation learning
computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Joint-Embedding Prediction
Compact World Representation
Efficient Training
Robust Robot Manipulation