Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

📅 2025-11-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing 3D vision pretraining methods exhibit limited performance in robotic manipulation, primarily due to the absence of explicit state-action-state dynamic modeling and excessive reliance on redundant explicit geometric reconstruction. This paper proposes AFRO, the first framework to formulate state prediction as a diffusion-based generative process, jointly learning forward and inverse dynamics within a shared latent space. AFRO introduces a feature-differencing mechanism and an inverse consistency constraint to suppress feature leakage in action representations and eliminate the need for geometric reconstruction. As a fully self-supervised approach, it significantly enhances the semantic discriminability and dynamic awareness of 3D visual representations. Evaluated on 16 simulated and 4 real-world manipulation tasks, AFRO substantially outperforms prior pretraining methods, achieving marked improvements in task success rates. Moreover, it demonstrates strong scalability with respect to both data volume and task complexity.

Technology Category

Computer Vision: Diffusion Models for VisionHumans and AI: Human-Aware Planning and Behavior PredictionIntelligent Robots: State Estimation

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We introduce AFRO, a self-supervised framework that learns dynamics-aware 3D representations without action or reconstruction supervision. AFRO casts state prediction as a generative diffusion process and jointly models forward and inverse dynamics in a shared latent space to capture causal transition structure. To prevent feature leakage in action learning, we employ feature differencing and inverse-consistency supervision, improving the quality and stability of visual features. When combined with Diffusion Policy, AFRO substantially increases manipulation success rates across 16 simulated and 4 real-world tasks, outperforming existing pre-training approaches. The framework also scales favorably with data volume and task complexity. Qualitative visualizations indicate that AFRO learns semantically rich, discriminative features, offering an effective pre-training solution for 3D representation learning in robotics. Project page: https://kolakivy.github.io/AFRO/
Problem

Research questions and friction points this paper is trying to address.

Learns dynamics-aware 3D representations without action or reconstruction supervision
Improves manipulation success rates across simulated and real-world robotic tasks
Scales effectively with increasing data volume and task complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative diffusion process for state prediction
Joint forward-inverse dynamics modeling in latent space
Feature differencing and inverse-consistency supervision
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qiwei Liang
Hong Kong University of Science and Technology (Guangzhou)
B
Boyang Cai
Hong Kong University of Science and Technology (Guangzhou)
M
Minghao Lai
Shenzhen University
S
Sitong Zhuang
Shenzhen University
T
Tao Lin
Beijing Jiaotong University
Yan Qin
Yan Qin
Chongqing University
Machine learningMultivariate statistical analysisAI
Yixuan Ye
Yixuan Ye
Data Scientist - Research, Google LLC
Statistical ModelingGenetic Prediction
J
Jiaming Liang
Hong Kong University of Science and Technology (Guangzhou)
Renjing Xu
Renjing Xu
HKUST(GZ)
Brain-inspired ComputingHumanoid Computing