joint human-scene reconstruction

Designs and implements methods that simultaneously reconstruct 3D human geometry (meshes, poses) and surrounding scene geometry into a shared metric coordinate frame, producing spatially consistent human placements within the environment. This competence covers building unified or single-pass models and training strategies that couple human and scene representations, share geometric features, and enforce alignment and physical plausibility between people and scenes.

jointhuman-scenereconstruction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of scale ambiguity, misalignment between humans and scenes, and occlusion interference in human reconstruction from dynamic scenes captured by a moving monocular camera. To tackle these issues, the authors propose SHOW—a framework that jointly infers human meshes and scene geometry in a unified metric space through a feed-forward pipeline. The method integrates the semantic structure and scale priors of parametric human models and employs a promptable mask mechanism to flexibly specify target individuals. Mutual guidance between human and scene representations enhances spatial alignment and metric scale consistency. Experiments demonstrate that SHOW significantly improves metric-scale reconstruction accuracy, human-scene alignment quality, and overall robustness in complex scenarios involving multiple people, severe occlusions, and cluttered backgrounds.

dynamic sceneshuman-scene misalignmentmonocular reconstruction

Reconstructing People, Places, and Cameras

Dec 23, 2024
LM
Lea Muller
🏛️ UC Berkeley

This work addresses the joint reconstruction of multi-person 3D meshes, scene point clouds, and camera poses from sparse, uncalibrated multi-view images, all within a unified metric world coordinate system that explicitly models spatial relationships among humans, environment, and cameras. Methodologically, it is the first to embed the SMPL human statistical model into a Structure-from-Motion (SfM) framework, leveraging human priors to impose absolute scale constraints. A multi-module joint optimization scheme is introduced to co-estimate human meshes, scene geometry, and camera parameters, synergistically integrating data-driven reconstruction with classical SfM principles. Evaluated on EgoHumans and EgoExo4D, the method reduces world-coordinate human localization error to 1.04 m and 0.56 m, respectively, and improves camera pose accuracy (RRA@15) by 20.3%, significantly enhancing overall geometric fidelity and cross-modal consistency.

Estimate metric scale using human statistical model for accurate scene and camera reconstruction.Improve human localization accuracy and camera pose estimation in world coordinate systems.Jointly reconstruct human meshes, scene point clouds, and camera parameters from uncalibrated multi-view images.

This work addresses the domain shift, geometric distortions, and misalignment commonly encountered when jointly reconstructing real-world scenes and 3D human structures due to reliance on synthetic data. The authors propose an end-to-end feedforward framework that simultaneously recovers metric-scale scene geometry, human point clouds, camera parameters, and SMPL body models in a single forward pass. Leveraging unlabeled real-world videos, the method innovatively integrates scene reconstruction with HMR priors through a two-stage training strategy—coarse localization using synthetic data followed by geometric refinement on real data—and employs a high-frequency detail distillation scheme to effectively bridge the sim-to-real domain gap. Experiments demonstrate that the approach achieves state-of-the-art performance in human-centric scene reconstruction and outperforms both optimization-based and pure HMR methods in global human motion estimation.

3D reconstructiondomain generalizationhuman-scene reconstruction

Physically Compatible 3D Object Modeling from a Single Image

May 30, 2024
MG
Minghao Guo
🏛️ MIT | UMass Amherst

This work addresses the lack of physical plausibility in single-image 3D reconstruction. We propose the first physics-compatible reconstruction framework that enforces static equilibrium as a hard constraint. Methodologically, we explicitly decouple and jointly optimize material stiffness, external loading forces, and the static equilibrium geometry; deformation responses are modeled via differentiable physics simulation, enabling gradient-based joint optimization of all variables. Our approach breaks from conventional simplifications—such as rigid-body assumptions or neglect of external forces—by embedding real-world physical constraints directly into the single-image reconstruction pipeline. Evaluated on Objaverse, our method yields reconstructions with significantly improved mechanical stability, suitable for downstream dynamic simulation and 3D printing. Physical validation via real-world force testing further confirms the structural robustness of the generated models.

3D modelingmaterial propertiesphysical stability

Human as Points: Explicit Point-based 3D Human Reconstruction from Single-view RGB Images

Nov 06, 2023
YT
Yingzhi Tang
🏛️ City University of Hong Kong | Tsinghua University

Existing implicit-representation-based methods for 3D human reconstruction from a single RGB image suffer from limited generalization, robustness, and geometric controllability. This paper proposes the first fully explicit, point-cloud-driven end-to-end framework—eliminating implicit functions entirely and directly regressing, generating, and optimizing human point clouds in 3D space. Key contributions include: (1) a geometric center regression paradigm to enhance structural controllability; (2) an SMPL-guided explicit point cloud estimation and refinement network; and (3) a prior-constrained point cloud generation and registration mechanism. Our method achieves significant improvements over state-of-the-art approaches across multiple benchmarks, with 20–40% reductions in Chamfer Distance (CD) and Mean Per-Joint Position Error (MPJPE), yielding more detailed and complete reconstructions. Code and data are publicly released.

3D reconstructiondeep learning limitationshuman model

Latest Papers

What's happening recently
View more

Existing multi-person 3D reconstruction methods often suffer from geometric incompleteness and implausible interpenetrations when handling occlusions and close interactions. This work proposes HUG3D, a novel framework that explicitly incorporates group-level contextual cues and physical interaction priors to jointly model individual and collective information in an orthonormal canonical space, enabling collaborative optimization of occlusion, contact, and spatial relationships. HUG3D comprises two core modules: Human Group-Instance Multi-View Diffusion (HUG-MVD) and Geometric Reconstruction (HUG-GR), which integrate multi-view normal generation with physics-aware geometric refinement to achieve high-fidelity texture fusion. Requiring only a single input image, HUG3D significantly outperforms current state-of-the-art single- and multi-person reconstruction approaches in both reconstruction fidelity and physical plausibility.

3D reconstructionmulti-human interactionocclusion handling

From Camera to World: A Plug-and-Play Module for Human Mesh Transformation

Dec 17, 2025
CM
Changhai Ma
🏛️ University of Science and Technology of China

3D human mesh reconstruction from in-the-wild images suffers from inaccurate orientation estimation in the world coordinate system, primarily due to the absence of ground-truth camera rotation—especially pitch angle—leading to substantial errors under the common zero-rotation assumption. To address this, we propose a human-centered strategy that estimates camera pitch solely from RGB images and synthetic depth maps. We further introduce a plug-and-play Mesh-Plug module that jointly optimizes root joint orientation and full-body pose. Additionally, we design a camera rotation prediction network grounded in human spatial configuration. Our method achieves significant improvements over state-of-the-art approaches on the SPEC-SYN and SPEC-MTP benchmarks, enabling more accurate and robust world-coordinate human mesh reconstruction without requiring real camera calibration.

Estimates camera rotation for 3D human mesh reconstructionRefines root joint orientation and body pose simultaneouslyTransforms human meshes from camera to world coordinates

Existing single-image human-scene interaction reconstruction methods struggle to balance speed and physical plausibility: optimization-based approaches are accurate but slow, while feedforward methods are fast yet lack explicit interaction modeling, often producing floating or interpenetration artifacts. This work proposes GRAFT, a learnable interaction prior that generates volume-anchored tokens via geometric probes and employs a lightweight recurrent Transformer to predict interaction gradients, thereby transforming geometric fitting into efficient feedforward inference. GRAFT achieves explicit and physically reasonable 3D interaction reconstruction at high speed. Experiments show that GRAFT improves interaction quality by up to 113% over state-of-the-art feedforward methods, operates approximately 50 times faster than optimization-based approaches, generalizes well to in-the-wild multi-person scenes, and achieves a user preference rate of 64.8%.

3D reconstructioncontact reasoninghuman-scene interaction

Existing methods for recovering 3D human geometry from monocular video face significant challenges: Vision Transformers (ViTs) tend to overfit to 2D viewpoints, while NeRF- and Gaussian Splatting–based avatars decouple pose and appearance, limiting generalization to novel poses. This work proposes HumanSplatHMR, the first approach to tightly couple human mesh recovery with Gaussian Splatting avatar modeling within a joint optimization framework. By leveraging differentiable rendering, photometric, segmentation, and depth losses are end-to-end backpropagated to pose parameters, enabling closed-loop optimization of both pose and appearance. Requiring neither motion capture nor offline refinement, the method substantially improves 3D pose accuracy in real-world scenes and enhances rendering fidelity under novel poses and viewpoints, outperforming existing decoupled baselines.

3D human poseavatar reconstructionGaussian Splatting

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
TC

Timothy Chen

Stanford University
RoboticsPerceptionControl
DR

Deva Ramanan

Professor, Robotics Institute, Carnegie Mellon University
Computer VisionMachine Learning
MS

Mac Schwager

Stanford University
RoboticsControlMulti-Agent SystemsMachine Learning
KD

Kostas Daniilidis

Ruth Yalom Stone Professor of Computer and Information Science, University of Pennsylvania
Computer VisionRobotics