Score
Designs and builds datasets for robotic systems by collecting, synchronizing, cleaning, annotating, and validating sensor, actuator, and metadata streams and associated labels or rewards. Implements data pipelines and formatters (e.g., RLDS-style formatting), serialization, validation, and versioning tools that convert raw robot recordings into standardized, trainable dataset artifacts and interfaces for training and evaluation.
This work addresses two core challenges in large-scale robotic manipulation datasets: (1) designing high-value diversity dimensions to enhance data utility, and (2) efficiently retrieving task-aligned demonstrations from existing datasets. To this end, we introduce a programmable data generation framework that explicitly models controllable diversity variables—including camera pose, object categories, and spatial layout. Our analysis reveals, for the first time, that camera pose and spatial arrangement are critical determinants of both dataset diversity and task alignment. We further propose a task-oriented demonstration retrieval algorithm grounded in geometric-semantic joint alignment. Evaluated on real-world datasets including DROID, our method improves downstream policy performance by up to 70%. Crucially, insights and gains observed in simulation generalize successfully to physical robot platforms, demonstrating robust cross-domain transferability.
Converting ROS bags into machine learning datasets often relies on ad hoc scripts, resulting in substantial engineering overhead and inefficient iteration. This work introduces, for the first time, the principles of software build systems to robotic dataset construction, proposing a reproducible, incremental generation method grounded in artifact- and dependency-graph semantics. We present Bagzel, an open-source tool built on Bazel, which supports export to the nuScenes format and incorporates Bagzel-xattr for server-side metadata management. Experimental evaluation demonstrates that, on a 20.4 GB dataset, hot builds achieve up to a 386.26× speedup and incremental builds are accelerated by 7.21×, with performance gains further amplified as dataset scale increases.
To address the challenges of large-scale, slow-loading, and hard-to-generalize multimodal robotic trajectory data (video, text, numerical), this paper proposes a cloud-native trajectory data management framework. We design EBML—a self-contained binary format supporting hybrid lossy/lossless compression—achieving up to 70× compression over RLDS without sacrificing downstream task accuracy. We introduce a novel memory-mapped decoding cache coupled with load-balanced, multi-stream parallel video decoding, accelerating decoding by 50× versus LeRobot. The full system maintains model performance even under 75× aggressive compression. This work establishes an efficient, scalable data infrastructure for training large-scale Transformer models across diverse robots and tasks.
Existing robot manipulation policies suffer from limited generalization due to reliance on small-scale, low-diversity simulation data or single-environment real-world datasets. To address this, we introduce DROID—the first large-scale, cross-household, real-world distributed robot manipulation dataset. It encompasses 564 diverse household environments, 84 task categories, and 76k high-quality trajectories (350 hours), collected over 12 months by 50 geographically distributed contributors. DROID pioneers intercontinental, multi-brand robotic hardware coordination (UR5e and Franka Emika arms) via remote distributed data collection, integrating standardized interfaces, precise action alignment, and rigorous quality filtering. We fully open-source the hardware specifications, data collection infrastructure, and training code. Policies trained on DROID achieve a 27% average success rate improvement in cross-scene generalization benchmarks and demonstrate superior zero-shot transfer performance compared to prior state-of-the-art methods.
Robot learning suffers from fragmentation, requiring separate model training for each robot platform, task, and environment. Method: This paper introduces X-robot, a universal policy paradigm enabling the first cross-platform positive transfer. We construct the largest standardized robotic manipulation dataset to date, developed collaboratively across multiple institutions. We propose RT-X, a Transformer-based architecture that unifies action representations and cross-robot data formats, trained via large-scale multi-robot behavioral cloning and transfer learning. Contribution/Results: RT-X demonstrates significant skill generalization across 22 heterogeneous robot platforms. In zero-shot and few-shot settings on unseen robots, it achieves substantial improvements in task success rates. This work provides the first systematic empirical validation of both the effectiveness and scalability of cross-robot positive transfer—establishing a foundational step toward generalizable robotic policies.
This work addresses the limited reproducibility of behavioral validation in robotic simulation testing, which often stems from insufficiently documented test configurations, execution protocols, and post-processing procedures. To overcome this, the study proposes a deep integration of data provenance and FAIR (Findable, Accessible, Interoperable, Reusable) principles throughout the entire test generation pipeline—rather than merely appending them to final datasets. The authors extend an existing simulation testing framework by embedding machine-readable, structured metadata at every stage, thereby enabling end-to-end traceable validation workflows. This approach significantly enhances the reproducibility of mobile robot navigation datasets. Additionally, the project distills practical FAIR implementation guidelines tailored to robotics, identifying key challenges such as vocabulary alignment, attribute selection, and adoption of community standards, and offers actionable recommendations for addressing them.
This work addresses critical bottlenecks in humanoid robotics—namely data silos, high acquisition costs, and inconsistent evaluation—stemming from the absence of a unified physical interaction data infrastructure, which hinders the scalable advancement of Physical AI. The study introduces the concept of “embodied interaction data” to characterize humanoid robot data and proposes a hierarchical architecture that integrates general standards with capability-specific ones. It emphasizes preserving the complete relational structure among robot embodiment, actions, tasks, environments, and outcomes, while ensuring physical consistency across multimodal data in terms of temporal alignment, coordinate frames, and calibration. Building upon the ISO/WD 26264-1 draft, the authors establish a horizontal data infrastructure encompassing metadata, provenance, quality control, and versioning, alongside domain-specific semantic specifications for manipulation, locomotion, and human–robot interaction. This framework provides an interpretable, shareable, and reusable data foundation to enable cross-platform, cross-task, and cross-institutional co-evolution in Physical AI.
This work addresses the lack of scalable, trustworthy evaluation methods and physically plausible training data for general-purpose robotic policies, compounded by the high cost and poor reproducibility of real-world robot experiments. To overcome these challenges, the authors propose a human↔simulation↔robot bidirectional alignment framework supported by a cloud-native toolchain. Leveraging the JoySim simulator—integrated with reconstruction, rendering, and realism-enhancement modules—they implement a high-fidelity digital twin on the JD Cloud platform. Human demonstrations are transformed into physically consistent trajectories, annotations, and visual observations, while simulation serves dual roles as a scalable evaluation layer and a data filter. This approach substantially improves both the efficiency of data generation and the reliability of policy evaluation.
This work investigates the sim-to-real transferability of vision-language-action (VLA) models to a physical UR5e robotic arm and presents an end-to-end pipeline encompassing real-world data collection, construction of an RLDS-compatible dataset, fine-tuning of OpenVLA models, and their deployment. By incorporating multimodal temporal alignment, unified coordinate systems, and consistent action semantics, the study demonstrates that successful real-world performance hinges on holistic co-design across the data–model–control pipeline rather than model performance alone. A reproducible evaluation framework reveals a substantial gap between offline metrics and closed-loop task success, primarily attributable to system-level factors such as action semantics, image preprocessing, and data quality.