🤖 AI Summary
This study addresses the geometric distortion and prediction incompleteness in monocular semantic scene completion for quadruped robots operating in crowded indoor environments, where human occlusion poses significant challenges. To this end, we propose the CrowdOcc dataset and an accompanying framework. Methodologically, a static-dynamic decoupled annotation strategy is introduced to eliminate occlusion ambiguity, while a normal-guided geometric fusion module is designed to enhance structural consistency. Furthermore, a human-centric sparse interaction mechanism is incorporated to strengthen human-robot relational modeling. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on unseen-scene generalization benchmarks, yielding an IoU of 15.80, an mIoU of 11.40, and a human IoU of 46.23. These findings indicate that our method significantly improves the robustness of 3D semantic understanding under complex occlusion scenarios.
📝 Abstract
Monocular semantic scene completion (SSC) for quadruped robots remains underexplored in real crowded indoor environments, where human-scene occlusion disrupts static geometry and human occupancy predictions are often incomplete or spatially misplaced. We present CrowdOcc, an RGB-D dataset and monocular SSC framework for this setting. CrowdOcc contains 25.1K frames from 11 indoor scenes, with semantic occupancy annotations constructed through static dynamic decoupling. Our framework combines: (i) Normal Guided Scene Geometry Fusion (NGSGF) to complement depth-aware lifting with surface-normal cues for occlusion robust geometry; and (ii) Human-Centric Sparse Interaction (HCSI) to selectively model human-human and local human scene relations in 3D. Our method achieves state-of-the-art SSC performance on CrowdOcc's scene-disjoint test set, reaching 15.80 IoU, 11.40 mIoU, and 46.23 Human IoU, demonstrating generalization to unseen indoor scenes.