Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes the first end-to-end embodied AI system fully built upon the AMD ROCm ecosystem, addressing the prevailing reliance of vision-language-action (VLA) models on CUDA. By integrating the SmolVLA model, 3D Gaussian Splatting, and the Genesis physics engine, the system enables full-stack acceleration—from data-center training and photorealistic simulation rendering to edge inference on Ryzen AI platforms—within the ROCm+PyTorch framework. It establishes a complete Real2Sim2Real loop and demonstrates successful deployment of language-guided semantic manipulation policies on a Franka robotic arm. Furthermore, large-scale reinforcement learning is validated across quadrupedal and humanoid robot platforms, confirming the feasibility and efficiency of a pure AMD software stack for complex embodied intelligence tasks.
📝 Abstract
Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the next big thing is Physical AI, AI with a body,'' GTC Paris, June 2025) and Dr. Lisa Su (`we're entering the world of Physical AI ... this is where AI enters the real world,' CES 2026). This paper presents an end-to-end, fully AMD-accelerated technology stack for embodied manipulation, spanning data-center training silicon, Radeon PRO simulation/rendering GPUs, and Ryzen AI edge compute, unified by the open ROCm software stack. We demonstrate that training and deploying VLA-based manipulation policies does not require a CUDA-locked ecosystem. Four progressive demonstrations are presented: (1) a Sim-to-Real manipulation pipeline trained with SmolVLA and deployed on a physical Franka arm; (2) a semantic, language-grounded object-selection task (`one-of-three'); (3) a Real2Sim synthetic-data generation pipeline that fuses 3D Gaussian Splatting (3DGS) reconstructions of real scenes with the Genesis physics engine; and (4) large-scale reinforcement learning for quadruped and humanoid locomotion benchmarked across multiple hardware platforms. All pipelines run natively on ROCm + PyTorch on RDNA4 (Radeon AI PRO R9700) and RDNA3.5 (Radeon PRO W7900) hardware and are reproducible on the free Radeon Cloud Platform.
Problem

Research questions and friction points this paper is trying to address.

Physical AI
Vision-Language-Action
Sim2Real
ROCm
Embodied Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Real2Sim2Real
Vision-Language-Action (VLA)
ROCm
3D Gaussian Splatting
Physical AI
🔎 Similar Papers