🤖 AI Summary
Existing 3D vision pretraining methods exhibit limited performance in robotic manipulation, primarily due to the absence of explicit state-action-state dynamic modeling and excessive reliance on redundant explicit geometric reconstruction. This paper proposes AFRO, the first framework to formulate state prediction as a diffusion-based generative process, jointly learning forward and inverse dynamics within a shared latent space. AFRO introduces a feature-differencing mechanism and an inverse consistency constraint to suppress feature leakage in action representations and eliminate the need for geometric reconstruction. As a fully self-supervised approach, it significantly enhances the semantic discriminability and dynamic awareness of 3D visual representations. Evaluated on 16 simulated and 4 real-world manipulation tasks, AFRO substantially outperforms prior pretraining methods, achieving marked improvements in task success rates. Moreover, it demonstrates strong scalability with respect to both data volume and task complexity.
📝 Abstract
Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We introduce AFRO, a self-supervised framework that learns dynamics-aware 3D representations without action or reconstruction supervision. AFRO casts state prediction as a generative diffusion process and jointly models forward and inverse dynamics in a shared latent space to capture causal transition structure. To prevent feature leakage in action learning, we employ feature differencing and inverse-consistency supervision, improving the quality and stability of visual features. When combined with Diffusion Policy, AFRO substantially increases manipulation success rates across 16 simulated and 4 real-world tasks, outperforming existing pre-training approaches. The framework also scales favorably with data volume and task complexity. Qualitative visualizations indicate that AFRO learns semantically rich, discriminative features, offering an effective pre-training solution for 3D representation learning in robotics. Project page: https://kolakivy.github.io/AFRO/