🤖 AI Summary
This work addresses the scarcity of real-world articulated object manipulation data and the lack of semantic and physical constraints in existing synthetic datasets by proposing the SMART system. Methodologically, it introduces SMART-Sim, a simulation platform featuring joint-aware design alongside a scalable distributed synthetic pipeline. Leveraging autonomous agents, the system automatically generates large-scale, high-quality demonstration data encompassing extensive atomic tasks and object categories, which is subsequently used to pre-train a Vision-Language-Action (VLA) model. Experimental results demonstrate that SMART achieves superior performance on simulation benchmarks. Furthermore, it enables zero-shot sim-to-real transfer and exhibits scalable performance improvements for articulated object manipulation in real-world settings.
📝 Abstract
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.