A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决VLA模型训练数据收集难题,提出了一种从模拟到现实的集成管道方法,通过在模拟中生成专家轨迹并在真实环境中重放来收集训练数据。
📝 Abstract
Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as sequences of action chunks. Training and evaluating these models requires large-scale collections of real-world demonstrations, pairing robot actions with the corresponding visual and proprioceptive observations. Collecting such data on real hardware typically relies on human teleoperation, making the process costly, time-consuming, and difficult to scale. We present an open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training. The same deployment stack is then reused, in closed-loop, to evaluate a trained policy on that setup, so that data collection and evaluation share an identical hardware configuration. Because each real recording is paired with the simulated trajectory that produced it, the protocol also yields a direct measurement of the sim-to-real gap. We release the collected datasets on Hugging Face together with the pipeline source code https://gitlab.isir.upmc.fr/kappel/sim2real_public_chunk_control.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
sim-to-real
data collection
robot actions
multimodal inputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sim-to-Real
VLA Models
Chunk-Based Policies
Data Collection
Policy Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mathilde Kappel
Institute of Intelligent Systems and Robotics (ISIR), CNRS, Sorbonne University, Paris, F-75005, France
C
Clémence Grislain
Institute of Intelligent Systems and Robotics (ISIR), CNRS, Sorbonne University, Paris, F-75005, France
Mohamed Chetouani
Mohamed Chetouani
Professor, Sorbonne Universite, ISIR-UPMC, CNRS
social signal processinghuman-robot interactioninteractive machine learningmultimodal interaction
Olivier Sigaud
Olivier Sigaud
Professor in Computer Science, Sorbonne Université
deep reinforcement learningartificial intelligencemachine learningdevelopmental roboticscomputational neuroscience
Louis Annabi
Louis Annabi
ISIR, Sorbonne Université
F
Faïz Ben Amar
Institute of Intelligent Systems and Robotics (ISIR), CNRS, Sorbonne University, Paris, F-75005, France
S
Stéphane Doncieux
Institute of Intelligent Systems and Robotics (ISIR), CNRS, Sorbonne University, Paris, F-75005, France
Mahdi Khoramshahi
Mahdi Khoramshahi
Associate Professor (MCF) at Sorbonne Université
RoboticsCollaborative RobotspHRIControl