Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Monocular 4D human crowd reconstruction in large-scale scenes is prone to depth ambiguity and complex terrain, often resulting in scale distortion and spatial drift. To address these challenges, this work proposes Crowd4D, the first scene-aware framework for monocular 4D crowd reconstruction. It introduces a Human-Scene Interaction Proxy (HSIP) as a unified optimization representation, integrates scene geometric priors, and incorporates a Crowd Structural Temporal Consistency Regularization (CSCR) to enhance stability under occlusions. By jointly optimizing both crowd and scene geometry, Crowd4D significantly outperforms existing methods on real-world large-scale sequences, achieving robust and scene-consistent 4D reconstructions.
📝 Abstract
Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex scene geometry. Existing monocular crowd reconstruction methods typically rely on single-plane assumptions, leading to unreliable metric scale and spatial drift under complex terrain. We propose Crowd4D, the first scene-aware 4D crowd reconstruction framework that jointly optimizes the crowd and scene from a monocular RGB video in large-scale scenes. Crowd4D explicitly incorporates scene geometry and ensures consistency across image and scene spaces via a multi-stage optimization strategy. A key bottleneck of this task lies in accurate human-scene alignment, particularly in scale and position. However, human and scene reconstructions are typically decoupled. To address this, we introduce the Human-Scene Interaction Proxy, abbreviated as HSIP, as an intermediate representation derived from Scene Interaction Point Clouds and a Scene Interaction Surface, abbreviated as SIPC and SIS. These representations encode explicit scene-aware geometric priors and redefine the optimization space for large-scale monocular 4D crowd reconstruction. To further improve temporal stability under occlusions, we introduce Crowd Structural Coherence Regularization, abbreviated as CSCR, which leverages HSIP-based spatial priors to impose soft temporal consistency on pairwise relative displacements and directions within local crowd neighborhoods. Extensive experiments demonstrate that Crowd4D consistently outperforms existing state-of-the-art methods and enables robust monocular 4D crowd reconstruction in complex, large-scale real-world scenes.
Problem

Research questions and friction points this paper is trying to address.

monocular crowd reconstruction
4D reconstruction
scene consistency
depth ambiguity
human-scene alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

scene-aware reconstruction
monocular 4D crowd
human-scene interaction proxy
structural coherence regularization
multi-stage optimization
🔎 Similar Papers
No similar papers found.