SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation
This study addresses the reliance of Vision-Language-Action (VLA) models on costly real-world data and their limited generalization in contact-rich manipulation tasks by proposing the SARI framework. Motivated by the insight that spatial coverage can be simulated while physical contacts require real-world grounding, SARI decouples tasks into two phases: simulated approach and real interaction. It leverages digital twins to generate diverse spatial trajectories and trains a unified policy with minimal real contact data. Seamless transfer without explicit labels is achieved through visual appearance alignment, shared camera-relative action representations, and co-training. Experimental results demonstrate that this approach reduces data collection time by 34.3% and achieves a 27.5% success rate on unseen object poses, significantly outperforming baselines.