Score
Design and implement algorithms and systems that recover 3D meshes (often human body meshes) expressed in the camera coordinate frame, producing vertex positions and poses in camera space and ensuring scene-consistent placement. This includes estimating or refining camera parameters and absolute scale so outputs are metric and localized in camera coordinates, and building evaluation/learning pipelines that work with or without known intrinsics to improve camera-space detection and localization accuracy.
This work addresses the joint reconstruction of multi-person 3D meshes, scene point clouds, and camera poses from sparse, uncalibrated multi-view images, all within a unified metric world coordinate system that explicitly models spatial relationships among humans, environment, and cameras. Methodologically, it is the first to embed the SMPL human statistical model into a Structure-from-Motion (SfM) framework, leveraging human priors to impose absolute scale constraints. A multi-module joint optimization scheme is introduced to co-estimate human meshes, scene geometry, and camera parameters, synergistically integrating data-driven reconstruction with classical SfM principles. Evaluated on EgoHumans and EgoExo4D, the method reduces world-coordinate human localization error to 1.04 m and 0.56 m, respectively, and improves camera pose accuracy (RRA@15) by 20.3%, significantly enhancing overall geometric fidelity and cross-modal consistency.
This work addresses the challenges of scale ambiguity, misalignment between humans and scenes, and occlusion interference in human reconstruction from dynamic scenes captured by a moving monocular camera. To tackle these issues, the authors propose SHOW—a framework that jointly infers human meshes and scene geometry in a unified metric space through a feed-forward pipeline. The method integrates the semantic structure and scale priors of parametric human models and employs a promptable mask mechanism to flexibly specify target individuals. Mutual guidance between human and scene representations enhances spatial alignment and metric scale consistency. Experiments demonstrate that SHOW significantly improves metric-scale reconstruction accuracy, human-scene alignment quality, and overall robustness in complex scenarios involving multiple people, severe occlusions, and cluttered backgrounds.
3D human mesh reconstruction from in-the-wild images suffers from inaccurate orientation estimation in the world coordinate system, primarily due to the absence of ground-truth camera rotation—especially pitch angle—leading to substantial errors under the common zero-rotation assumption. To address this, we propose a human-centered strategy that estimates camera pitch solely from RGB images and synthetic depth maps. We further introduce a plug-and-play Mesh-Plug module that jointly optimizes root joint orientation and full-body pose. Additionally, we design a camera rotation prediction network grounded in human spatial configuration. Our method achieves significant improvements over state-of-the-art approaches on the SPEC-SYN and SPEC-MTP benchmarks, enabling more accurate and robust world-coordinate human mesh reconstruction without requiring real camera calibration.
This work addresses the challenging problem of joint 3D human mesh reconstruction and camera pose estimation from uncalibrated multi-view videos—particularly under severe occlusion and textureless conditions. We propose an end-to-end differentiable framework that jointly optimizes human body meshes and camera parameters by introducing two novel components: (i) a pose-geometry consistency constraint enforcing geometric coherence across views and poses, and (ii) an implicit motion prior encoding natural human dynamics. Initialization leverages 2D human detections and keypoint heatmaps to estimate intrinsic and extrinsic camera parameters; subsequent optimization integrates cross-view person association, differentiable rendering, and motion modeling—requiring neither calibration targets nor background features. Evaluated on multiple public benchmarks, our method achieves single-stage reconstruction with 37% lower camera parameter error and an MPJPE of 58.3 mm, significantly outperforming conventional sequential approaches. The code is publicly available.
This work proposes a ray-based camera calibration framework tailored for 3D reconstruction, addressing the limitations of traditional reprojection error–based methods that rely on 2D calibration boards and inadequately reflect 3D geometric accuracy. Instead of reprojection error, the approach introduces reconstruction error and intersection error as more representative metrics. It employs a novel icosahedral 3D calibration target and a ring-shaped feature detector, integrated with a generalized distortion model and bootstrapping to refine both intrinsic and extrinsic parameter estimates. Experimental results on synthetic data demonstrate that the proposed method reduces average intersection error by approximately 40%, significantly enhances calibration stability, and validates that ray-level metrics provide a more faithful assessment of 3D reconstruction fidelity compared to conventional approaches.
This work addresses the ambiguity in distal joint poses during human mesh recovery caused by occlusion or insufficient depth cues. The authors propose a two-stage hybrid framework that first employs a deterministic regression module to stably recover the torso and root anchor points, followed by probabilistic flow matching to generate plausible poses for non-torso body parts. The method innovatively differentiates the treatment of torso and distal joints by integrating composite target representations, geometry-aware supervision, and feature-aware classifier-free guidance. Additionally, a multi-view synthetic data pipeline is introduced to provide image–camera–motion paired supervision. The approach achieves state-of-the-art performance on both camera-space and world-space benchmarks, demonstrating significant improvements over strong baselines—particularly under heavy occlusion and on world-space metrics sensitive to drift.
Egocentric full-body mesh reconstruction from monocular head-mounted cameras struggles to recover fine-grained details of hands and face due to the lack of reliable ground-truth annotations based on parametric models like SMPL-X. This work proposes a prior-guided learning framework that generates high-fidelity pseudo ground truth through optimization-driven refinement. By integrating exocentric HMR model transfer, a diffusion-based pose prior, and a deterministic fisheye undistortion module, our approach achieves, for the first time, high-fidelity egocentric reconstruction of the entire body—including hands and face. Extensive evaluations demonstrate significant improvements over existing methods across multiple benchmarks, with the proposed pseudo ground truth outperforming conventional regression-based strategies. The code and annotated data are publicly released.
Existing methods for recovering 3D human geometry from monocular video face significant challenges: Vision Transformers (ViTs) tend to overfit to 2D viewpoints, while NeRF- and Gaussian Splatting–based avatars decouple pose and appearance, limiting generalization to novel poses. This work proposes HumanSplatHMR, the first approach to tightly couple human mesh recovery with Gaussian Splatting avatar modeling within a joint optimization framework. By leveraging differentiable rendering, photometric, segmentation, and depth losses are end-to-end backpropagated to pose parameters, enabling closed-loop optimization of both pose and appearance. Requiring neither motion capture nor offline refinement, the method substantially improves 3D pose accuracy in real-world scenes and enhances rendering fidelity under novel poses and viewpoints, outperforming existing decoupled baselines.
Existing single-view mesh reconstruction methods exhibit poor generalization under camera rotation due to their reliance on viewpoint priors, often resulting in 3D inconsistencies, erroneous scene layouts, and violations of physical constraints. This work proposes the first evaluation protocol specifically designed for single-view reconstruction under camera rotation, enabling systematic assessment of depth estimation, object meshes, scene layout, and physical plausibility. A two-stage pipeline built upon SAM3D and FoundationPose—augmented with ICP registration, monocular depth estimation, and gravity alignment—significantly enhances robustness. Furthermore, a novel gravity-aware refinement strategy reduces layout orientation error by 47.1% compared to single-stage approaches.