🤖 AI Summary
Existing social navigation datasets lack sufficient diversity across cultural, geographical, and human–robot interaction dimensions, limiting their ability to support generalizable models of social behavior. To address this gap, this work presents a large-scale data collection effort deploying seven distinct robotic platforms across eight locations in five countries. For the first time, it systematically integrates multi-cultural contexts, diverse robot embodiments, and explicit verbal interactions to capture multimodal navigation data in complex social scenarios. The resulting dataset combines 3D/2D visual inputs, odometry, overhead pedestrian tracking, and manually annotated trajectories, yielding 29.35 hours of on-robot data and 43.5 hours of overhead tracking data. This effort substantially expands the data boundaries and behavioral coverage for social navigation research.
📝 Abstract
Understanding how robots and humans move in shared spaces is essential for designing effective social robot navigation policies and predicting human behavior. However, existing datasets often lack the diversity needed to capture differences in culture, geography, and human-robot interaction-factors that strongly shape appropriate social behavior. To address this gap, we introduce ACME: A Cross-cultural, Multi-Embodiment dataset for social navigation. A large-scale data collection effort across 8 sites in 5 countries, using 7 robot embodiments, ACME is a large and diverse multi-modal dataset aimed at advancing social navigation research, providing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data. Unlike prior datasets, it focuses on capturing goal-driven social navigation behavior in complex social scenarios with explicit robot-crowd interaction through robot speech. To facilitate learning navigation policies and predicting pedestrian trajectories, ACME provides 3D and 2D scene features, odometry, interaction information, and human-annotated pedestrian trajectory labels. We make ACME easy to use by providing both human-readable data for each sensor modality as well as raw binary data. Our qualitative and quantitative analyses show that our dataset captures more challenging scenarios and a broader distribution of pedestrian behavior than previous datasets.