SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity of real-world articulated object manipulation data and the lack of semantic and physical constraints in existing synthetic datasets by proposing the SMART system. Methodologically, it introduces SMART-Sim, a simulation platform featuring joint-aware design alongside a scalable distributed synthetic pipeline. Leveraging autonomous agents, the system automatically generates large-scale, high-quality demonstration data encompassing extensive atomic tasks and object categories, which is subsequently used to pre-train a Vision-Language-Action (VLA) model. Experimental results demonstrate that SMART achieves superior performance on simulation benchmarks. Furthermore, it enables zero-shot sim-to-real transfer and exhibits scalable performance improvements for articulated object manipulation in real-world settings.
📝 Abstract
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
Problem

Research questions and friction points this paper is trying to address.

Articulated Object Manipulation
Sim-to-Real Transfer
Synthetic Data
Embodied AI
Vision-Language-Action Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Articulated Object Manipulation
Sim-to-Real Transfer
Vision-Language-Action Model
Synthetic Pretraining
Agentic Task Generation
🔎 Similar Papers
2024-07-16Neural Information Processing SystemsCitations: 16
💼 Related Jobs
No related jobs found.
J
Jicong Ao
Institute of Artificial Intelligence, China Telecom
S
Shuhan Jiang
Institute of Artificial Intelligence, China Telecom; Zhejiang University
Y
Yuling Zhong
Institute of Artificial Intelligence, China Telecom; Technical University of Munich
Y
Yanwen Liu
Institute of Artificial Intelligence, China Telecom; Technical University of Munich
Y
Yuhan Gao
Institute of Artificial Intelligence, China Telecom; Harbin Institute of Technology
J
Jiangyuan Zhao
Institute of Artificial Intelligence, China Telecom; Shanghai Jiao Tong University
Y
Yang Zhang
Institute of Artificial Intelligence, China Telecom; Tsinghua University
S
Shiqiang Zhu
Zhejiang University
Chenjia Bai
Chenjia Bai
Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院, TeleAI)
Reinforcement LearningRoboticsEmbodied AI
X
Xuelong Li
Institute of Artificial Intelligence, China Telecom