🤖 AI Summary
This study addresses the underexploited value of simulated demonstrations in Sim-to-Real co-training, where unified metrics for trajectory utility and set coverage remain absent. To systematically investigate data curation in this context, we propose TUCO, a novel framework that leverages influence functions to decompose the contribution of source-domain demonstrations to the target domain. TUCO establishes a unified metric system grounded in closed-loop behavior and employs a performance-aligned subset optimizer to select complementary simulation data for enhanced policy training. Extensive experiments on the RoboMimic and OmniReset benchmarks demonstrate that proactive data curation significantly improves training efficiency and achieves state-of-the-art performance across diverse settings.
📝 Abstract
Simulation demonstrations can supplement scarce real-world data for robot policy co-training. However, the value of using data curation to actively select these demonstrations for sim-to-real co-training remains underexplored. Existing curation methods also lack a unified criterion for measuring trajectory-level utility and set-level coverage from closed-loop target behavior. To address these gaps, we present the first systematic study of data curation for sim-to-real robot policy co-training and propose Trajectory-level Utility and set-level Coverage Optimization (TUCO). TUCO uses influence functions to trace how each source demonstration affects target-domain scoring rollouts. Our key insight is that these effects can be decomposed into an overall contribution to target return and variation across rollouts, providing a common closed-loop basis for measuring trajectory utility and set coverage. We further propose a performance-aligned subset optimizer that combines these measures in a unified curation objective to reduce redundancy and select complementary demonstrations. Extensive experiments on RoboMimic and OmniReset establish the value of active simulation data curation for sim-to-real policy co-training and show that TUCO achieves state-of-the-art performance across single-simulator, sim-to-sim, and sim-to-real settings.