ECHO: Embodied Camera Observations of Human Object Carrying
This study addresses the lack of benchmarks for reasoning about natural object placement locations for embodied agents in complex environments. To this end, it introduces the contextual object placement task and constructs ECHO, the first large-scale synthetic dataset for this purpose. The proposed approach uniquely integrates RGB-D scene scans, human motion trajectories, and natural language context. By leveraging HM3D-based generation techniques, 6-DoF trajectory tracking, and input mask probes for multimodal evaluation, it establishes a novel benchmark requiring joint reasoning over scene structure and user habits. Experimental results demonstrate that no single modality suffices for this task, confirming the necessity of fusing scene structure, human activity, and contextual knowledge to achieve precise object placement.