ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
ARSTAG系统通过将单张RGB图像和自然语言指令转化为机器人学习数据,解决了适应新操作任务需大量手动工程或遥操作数据收集的问题。
📝 Abstract
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Problem

Research questions and friction points this paper is trying to address.

visuomotor policies
manipulation tasks
task-specific data
simulation
data generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Real2Sim2Real
natural-language instruction
task-consistent randomization
sim-to-real transfer
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
B
Bowei Li
Carnegie Mellon University, Pittsburgh, PA 15213, USA
Y
Yuner Zhang
Carnegie Mellon University, Pittsburgh, PA 15213, USA
Changliu Liu
Changliu Liu
Associate Professor, Carnegie Mellon University
Roboticshuman-robot interactionsmotion planningoptimizationmulti-agent system