SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of Vision-Language-Action (VLA) models on costly real-world data and their limited generalization in contact-rich manipulation tasks by proposing the SARI framework. Motivated by the insight that spatial coverage can be simulated while physical contacts require real-world grounding, SARI decouples tasks into two phases: simulated approach and real interaction. It leverages digital twins to generate diverse spatial trajectories and trains a unified policy with minimal real contact data. Seamless transfer without explicit labels is achieved through visual appearance alignment, shared camera-relative action representations, and co-training. Experimental results demonstrate that this approach reduces data collection time by 34.3% and achieves a 27.5% success rate on unseen object poses, significantly outperforming baselines.
📝 Abstract
Vision-language-action (VLA) models often require costly real-world demonstrations to adapt to contact-rich manipulation tasks, particularly when generalization across object placements is needed. We propose SARI (Simulated Approach, Real Interaction), a phase-split sim-and-real co-training framework built on a simple insight: spatial coverage and contact physics should be acquired from the domains best suited to them. Specifically, free-space approaches require spatial diversity but tolerate modest simulation gaps, making them ideal for synthetic generation; conversely, contact interactions demand accurate physics but vary little across object placements, allowing a few real demonstrations to generalize across the workspace. SARI generates diverse simulated approaches in a photorealistic digital twin while collecting real contact interactions at only a few placements. Post-trained on these phase-segmented demonstrations, a single policy seamlessly stitches simulated approaches with real contact interactions using visual appearance alignment and a shared camera-relative action representation--without explicit phase labels or hand-coded switches. Across five real-world contact-rich manipulation tasks, SARI reduces real-data collection time by 34.3% and achieves 27.5% success at unseen placements, where all full-task sim-real baselines fail completely (0%).
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Contact-Rich Manipulation
Sim-to-Real Transfer
Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phase-split sim-real co-training
Vision-language-action models
Contact-rich manipulation
Digital twin
Camera-relative action representation
💼 Related Jobs
No related jobs found.
X
Xingxin He
The Hong Kong University of Science and Technology
Y
Yuxuan Jiang
The Hong Kong University of Science and Technology
H
Haonan Zhang
The Hong Kong University of Science and Technology
C
Chuhan Cui
The Hong Kong University of Science and Technology
K
Kaile Li
The Hong Kong University of Science and Technology
Z
Zhongxing Zheng
BYD Company Limited
C
Caihao Xu
BYD Company Limited
Z
Ziqi Wang
The Hong Kong University of Science and Technology