AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

πŸ“… 2026-07-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing robotic manipulation learning data pipelines, which are often constrained by specialized hardware, centralized teleoperation, or fixed task sets, resulting in poor scalability and limited diversity. To overcome these challenges, the authors propose AXISβ€”a scalable, community-driven data engine that collects large-scale demonstrations via browser-based teleoperation and integrates automated task generation, success detection, trajectory smoothing, and multimodal augmentation to construct high-quality training data. AXIS introduces a novel community collaboration framework enabling automatic task expansion, data curation, augmentation, and systematic evaluation. The resulting dataset encompasses 207 tasks and over 50,000 trajectories. Continuous pretraining on AXIS improves the success rate of Ο€β‚€.β‚… by 5.8% and outperforms the RoboCasa365 model by 37.3%, demonstrating substantial generalization capabilities in perturbed environments.
πŸ“ Abstract
Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or fixed task suites. We present AXIS, a growable community-driven data engine and benchmark for scalable robot learning, which enables browser-based teleoperation for large-scale demonstration collection, automatically generates and validates new manipulation tasks, and transforms community-collected demonstrations into training-ready data through automated success checking, quality filtering, trajectory smoothing, and visual and physics-based augmentation. The AXIS dataset currently contains 207 diverse tasks and 50K+ trajectories. Meanwhile, AXIS organizes data into task snapshots and evaluates policies with a systematic held-out protocol. We compare vision-language-action (VLA) policies under a unified AXIS evaluation suite and analyze scaling behavior across different data volumes. Continual pretraining on AXIS substantially improves the overall success rate of $Ο€_{0.5}$ by 5.8%, outperforms the model pretrained on RoboCasa365 by 37.3%, and exhibits consistent scaling with increasing data volume, with the largest gains observed under layout, sensor-noise, and camera perturbations.
Problem

Research questions and friction points this paper is trying to address.

scalable robot manipulation
data pipeline
demonstration collection
task diversity
robot learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

community-driven data engine
scalable robot manipulation
automated task generation
trajectory augmentation
vision-language-action policies
πŸ”Ž Similar Papers
M
Mengfei Zhao
Axis Robotics
D
Dihong Huang
University of California, Berkeley
Y
Yikai Tang
University of California, Berkeley
Peihao Li
Peihao Li
Tsinghua University; Tongyi Lab, Alibaba
3D reconstructiongenerative modelhuman avatar
Mingxuan Yan
Mingxuan Yan
University of California, Riverside
R
Ruiqi Zhuang
Axis Robotics
Y
Yanjia Huang
Texas A&M University
J
Jie Wang
Axis Robotics, Johns Hopkins University
H
Hai Zhai
Axis Robotics
T
Tony Zhou
Axis Robotics, University of Pennsylvania
R
Rui Zhang
Axis Robotics, University of Michigan
Z
Zhexi Luo
Axis Robotics, National University of Singapore
Yuchen Huang
Yuchen Huang
University of Michigan - Ann Arbor
AI InterpretabilityMachine LearningNeural SystemsUbiquitous Computing
Jianfei Yang
Jianfei Yang
Assistant Professor, Director of MARS Lab, Nanyang Technological University
Physical AIEmbodied AIMultimodal AI
J
Jiachen Li
Georgia Institute of Technology