WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of annotations and absence of identity association in metric 4D human-scene reconstruction from monocular videos by proposing a feed-forward framework grounded in 3D foundation models. Methodologically, pseudo-label pre-training and lightweight adaptation enable high-fidelity metric reconstruction of cameras, scenes, and humans. The authors introduce intermediate-layer token probing and a dustbin-aware Sinkhorn matching mechanism to achieve metric scale and identity consistency without additional inference, alongside a Scale Readout module and analytical Sim(3) composition for alignment optimization. Experiments demonstrate that this approach is the first feed-forward method to surpass optimization-based approaches on EMDB-2 while outperforming all existing feed-forward baselines. It further exhibits strong performance on the RICH dataset, achieving real-time inference at 10.1 fps on a single GPU.
📝 Abstract
3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through lightweight adaptation. Exact metric labels are scarce, but unlabeled in-the-wild video is abundant. We use people in curated web video to initialise the solution: a posed metric body and 2D keypoints give an approximate, closed-form scale pseudo-label. These pseudo-labels pretrain a Scale Readout, which is then fine-tuned together with a lightweight adapter using exact metric supervision from standard real-video training splits. At inference the head predicts metric scale from foundation-model tokens, without the ruler or its teachers. For person identity, we probe the pretrained foundation model alone and find evidence that its intermediate query-key features encode person correspondence across frames. In most evaluated moving-person clips, a mid-layer token prefers that person over the vacated location and other people. A tiny projection reads this correspondence; together with metric pelvis motion and proposal confidence, it drives dustbin-aware Sinkhorn association of per-frame bodies. WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video. Each window is predicted feed-forward; analytic association and Sim(3) composition connect windows. On EMDB-2, WildHSR is the first feed-forward method in the published comparison to beat the best optimization-based WA-MPJPE and RTE while leading feed-forward methods on all three world-frame metrics. On RICH, it leads feed-forward people-and-scene methods on WA-MPJPE and W-MPJPE. The complete pipeline runs at 10.1 fps on one GPU.
Problem

Research questions and friction points this paper is trying to address.

4D people-scene reconstruction
metric scale
person identity
3D foundation model
monocular video
Innovation

Methods, ideas, or system contributions that make the work stand out.

metric scale recovery
person identity association
feed-forward reconstruction
3D foundation model adaptation
Sinkhorn matching
🔎 Similar Papers
No similar papers found.