ECHO: Embodied Camera Observations of Human Object Carrying

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of benchmarks for reasoning about natural object placement locations for embodied agents in complex environments. To this end, it introduces the contextual object placement task and constructs ECHO, the first large-scale synthetic dataset for this purpose. The proposed approach uniquely integrates RGB-D scene scans, human motion trajectories, and natural language context. By leveraging HM3D-based generation techniques, 6-DoF trajectory tracking, and input mask probes for multimodal evaluation, it establishes a novel benchmark requiring joint reasoning over scene structure and user habits. Experimental results demonstrate that no single modality suffices for this task, confirming the necessity of fusing scene structure, human activity, and contextual knowledge to achieve precise object placement.
📝 Abstract
Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it. Progress on this problem has been limited, in part because no dedicated benchmark or dataset exists to define and evaluate it. Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object's natural destination. We introduce contextual object placement as a benchmark task: predicting an object's destination during an observed object-carrying episode. To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations. ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations. It comprises 3,805 human-annotated episodes across 159 floors of 115 HM3D scenes, involving 198 distinct objects. Each floor includes a complete RGB-D scan with human-annotated room labels and a surface list. Each episode provides synchronized RGB-D encounter clips; 6-DoF camera, human, and object trajectories; start and destination surfaces; an action caption; and a human-written context: a single sentence describing the inhabitant's routine that implies the destination without naming it. We evaluate contextual object placement using input-masked probes and an end-to-end baseline. Results show that no single input modality is sufficient, highlighting the need to jointly reason over scene structure, human activity, and contextual knowledge.
Problem

Research questions and friction points this paper is trying to address.

contextual object placement
embodied agents
human-object interaction
destination prediction
benchmark dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contextual Object Placement
Embodied AI
RGB-D Dataset
Human-Object Interaction
Multimodal Reasoning
💼 Related Jobs
No related jobs found.
X
Xuefei Sun
Intelligent Robotics Laboratory, Department of Computer Science, University of Colorado Boulder, Boulder, CO 80309, USA
Lorin Achey
Lorin Achey
Ph.D. Student, University of Colorado Boulder
Robot PerceptionAutonomous VehiclesGenerative AIEmbodied AI
K
Kali Hamilton
Intelligent Robotics Laboratory, Department of Computer Science, University of Colorado Boulder, Boulder, CO 80309, USA
Alberto Speranzon
Alberto Speranzon
Chief Scientist | Lockheed Martin | Advanced Technology Labs
Multi-agent SystemsNetworked SystemsMachine LearningApplied Category Theory
G
Gregory Grebe
Lockheed Martin, Advanced Technology Laboratories, USA
Yonatan Bisk
Yonatan Bisk
Assistant Professor, Carnegie Mellon University
Natural Language ProcessingEmbodied AIRobot Learning
Christoffer Heckman
Christoffer Heckman
Associate Professor, University of Colorado
roboticsautonomyperception