ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Embodied intelligence is hindered by the scarcity of multimodal, perception-to-action closed-loop data that fully captures natural human behavior across diverse viewpoints, modalities, and spatial scales. To address this, this work introduces ACE, an environment-embedded capture system that synchronously fuses first-person and multi-view video, full-body and hand motion, 6-DoF object trajectories, audio, and tactile signals within real-world home settings. ACE enables, for the first time, cross-scale, highly synchronized, multimodal recording of behaviors ranging from fine-grained tabletop manipulation to room-scale whole-body activities, while preserving behavioral diversity under object-level task instructions. The resulting ACE-Data-0 dataset comprises 150 hours (17 million frames) of interactions across 200 tasks, 50 participants, and 75,000 interaction clips, revealing critical bottlenecks in current multimodal models regarding contact reasoning, occlusion handling, ego-motion, and long-horizon temporal modeling.
📝 Abstract
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
Problem

Research questions and friction points this paper is trying to address.

embodied intelligence
multimodal data
perception-action loop
human-centric data
data bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

embodied intelligence
multisensory synchronization
human-centric data engine
perception-action loop
scalable imitation learning
🔎 Similar Papers
No similar papers found.
Yukang Cao
Yukang Cao
Research Fellow, Nanyang Technological University
3D computer vision
Haozhe Xie
Haozhe Xie
Nanyang Technological University
Computer Vision3D VisionGenerative AIRobotics
B
Beichen Wen
ACE Robotics
R
Runmao Yao
S-Lab, Nanyang Technological University
Y
Yinghao Liu
ACE Robotics
Y
Yue Huang
ACE Robotics
Z
Zhichao Liao
ACE Robotics
Y
Yunxiang Wang
S-Lab, Nanyang Technological University
H
Haiheng Liu
S-Lab, Nanyang Technological University
X
Xingshun Tian
ACE Robotics
D
Dawei Su
ACE Robotics
L
Long Zhuo
ACE Robotics
Dacheng Tao
Dacheng Tao
Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processingdata mining
X
Xiaogang Wang
ACE Robotics
L
Liang Pan
ACE Robotics
Ziwei Liu
Ziwei Liu
Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics