Score
Designs and evaluates methods that convert relative depth predictions from a single-camera (monocular) estimator into metric-scale depth maps, including scale-and-shift correction models, scale-aware inference, and self-supervised fine-tuning procedures. Builds calibration algorithms, training losses, and validation metrics to reduce absolute depth error and produce reliable, metrically calibrated monocular depth for downstream perception or navigation tasks.
Monocular depth estimation (MDE) suffers from inconsistent scale due to the absence of metric ground truth, severely limiting its applicability in downstream tasks such as visual localization and 3D reconstruction. This work addresses zero-shot monocular metric depth estimation (MMDE), aiming to improve cross-scene generalization and boundary detail fidelity. We first systematically survey the evolution of scale-invariant approaches and establish zero-shot MMDE as the canonical benchmark paradigm. Our method innovatively integrates label-free data augmentation, image patch-based modeling, geometry-aware architectural optimization, and generative prior learning. Experiments demonstrate substantial improvements in depth scale consistency and structural integrity across multiple benchmarks. Under the zero-shot setting, our approach achieves superior calibration accuracy and generalization performance compared to prior methods. Furthermore, we identify key open challenges and promising directions for future research in metric monocular depth estimation.
Monocular depth estimation suffers from scale ambiguity and domain shift, making metric depth recovery challenging. This work proposes a language-guided uncertainty envelope mechanism that leverages textual descriptions to provide coarse-scale priors and adaptively selects image-specific affine calibration parameters within an uncertainty-aware range, thereby avoiding reliance on noisy text-point estimates. The approach freezes both the relative depth backbone and the CLIP text encoder, integrating multi-scale visual feature pooling with an inverse-depth-space affine transformation to enable efficient and lightweight calibration. It achieves improved in-domain accuracy on NYUv2 and KITTI and demonstrates significantly better zero-shot transfer performance on SUN-RGBD and DDAD compared to purely language-based baselines, exhibiting enhanced robustness.
Monocular depth prediction suffers from inherent scale and offset ambiguities, leading to inaccurate estimation of relative camera pose. To address this, we propose a unified optimization framework that jointly estimates depth scale, depth offset, and camera pose—formulating and solving this previously unaddressed trivariate coupling problem for the first time. For three calibration settings—fully calibrated, shared focal length, and independently estimated focal lengths—we design efficient closed-form analytic solvers that integrate point correspondences and monocular depth constraints. Our approach combines nonlinear least-squares optimization with extended PnP techniques tailored to multiple scenarios. We evaluate the method on synthetic data and two large-scale real-world benchmarks, covering 11 state-of-the-art monocular depth predictors. Results demonstrate superior robustness and achieve state-of-the-art localization accuracy across diverse settings.
This paper addresses key challenges in monocular depth estimation—dependency on camera intrinsics, absence of absolute scale, blurry high-frequency details, and low inference efficiency—by proposing a zero-shot, calibration-free, high-resolution metric depth estimation method. Methodologically, it introduces: (1) an efficient multi-scale vision transformer architecture tailored for dense prediction; (2) a joint real-and-synthetic data training paradigm with a boundary-aware loss function; and (3) an end-to-end single-image focal-length regression module enabling absolute-scale recovery without prior intrinsic parameters. The approach achieves state-of-the-art performance on NYUv2 and KITTI, with markedly improved depth map boundary sharpness. At 2.25 megapixels input resolution, inference takes only 0.3 seconds on a standard GPU. Code and pretrained models are publicly released.
To address the challenge of calibrating affine-invariant disparities to metric depth in zero-shot monocular depth estimation, this paper proposes a fine-tuning-free test-time adaptive rescaling method. Leveraging sparse 3D points—e.g., from low-resolution LiDAR, SfM, or IMU reconstructions—as geometric priors, the method decouples the affine ambiguity in predicted depth maps and jointly corrects scale and shift in real time. Its key contribution is the first zero-parameter-modification approach to metric depth recovery: it preserves the pretrained model’s strong generalization capability while maintaining robustness against both sparse input noise and model prediction errors. Evaluated on NYUv2 and KITTI, the method outperforms existing zero-shot approaches, matches the performance of fully supervised fine-tuning, and significantly surpasses depth completion-based methods.
Existing monocular depth estimation methods produce only relative depth without metric scale, while surface normal estimation suffers from poor generalization to unseen scenes. This paper introduces the first monocular geometric foundation model capable of zero-shot, image-level joint estimation of metric depth and surface normals, enabling plug-and-play 3D metric reconstruction for arbitrary camera parameters and unknown scenes. Key contributions include: (1) a camera-agnostic canonical camera space transformation module that explicitly decouples and resolves depth scale ambiguity; (2) a depth-normal joint optimization mechanism that enhances zero-shot generalization of normal estimation; and (3) training on a large-scale, heterogeneous dataset comprising 16 million images and 1,000 camera models with diverse annotations. Experiments demonstrate state-of-the-art performance on both metric depth and surface normal estimation, significantly mitigating scale drift in monocular SLAM and enabling high-fidelity, dense metric 3D mapping from Internet-sourced images.
Monocular geometric estimation often suffers from “scale collapse,” leading to severe underestimation of true scene scale, particularly in distant or large-scale environments. To address this, this work presents an improved variant of the MoGe-2 model. It first introduces MetricScenes, the first large-scale in-the-wild metric-scale dataset derived from internet-sourced photographs and stereo imagery, leveraging geotags and known baselines to recover absolute scale. The method further incorporates a two-stage Poisson depth completion strategy that refines depth quality using estimated camera poses and initial depth maps. This approach substantially mitigates scale collapse, achieving superior metric accuracy in open-domain scenes while maintaining state-of-the-art performance on standard benchmarks.
This work addresses the limited generalizability of existing monocular depth estimation methods across diverse camera types and the scarcity of densely annotated panoramic data. The authors propose DepthMaster, a framework that decomposes panoramic images into overlapping perspective patches, thereby eliminating geometric discrepancies through a unified perspective representation. By introducing a correspondence consistency loss (CCL) and leveraging virtual projection cameras as geometric priors, the method enables seamless depth stitching without modifying the backbone network. Employing a Transformer-based architecture and a hybrid training strategy, DepthMaster achieves zero-shot state-of-the-art performance across 13 diverse datasets using only a single panoramic dataset for training—marking the first demonstration of a unified model capable of metric depth estimation for both narrow-field and 360° panoramic imagery.
OmniPoint通过采用解耦射线和距离表示、双向增强策略及鲁棒信息注入机制,解决了单目图像重建3D几何时对相机模型假设的限制问题。
This work addresses the challenge of metrically inaccurate monocular depth estimation on non-Lambertian surfaces—such as transparent or specular objects—which hinders reliable robotic manipulation and navigation. The authors propose a training-free depth alignment framework that leverages factor graph optimization to locally align monocular depth priors with raw sensor depth via affine transformations, preserving geometric details and boundary discontinuities while achieving metric accuracy. Key contributions include the first training-free method for metric alignment of monocular depth, the introduction of the first dense ground-truth benchmark dataset encompassing full-scene non-Lambertian objects—overcoming reliance on synthetic CAD models—and a novel data collection strategy combining multi-camera fusion with matte reflective spray. Evaluated across diverse sensors and complex real-world scenes, the approach significantly improves depth accuracy without any training, and the code is publicly released.
Existing monocular metric depth estimation methods exhibit limited generalization across diverse camera types such as fisheye and 360° cameras, hindering unified and accurate depth prediction. This work proposes a decoupling strategy that separates the task into relative depth prediction and spatially varying scale estimation. We introduce a lightweight depth-guided scale estimation module and a distortion-aware positional encoding, RoPE-φ, which incorporates latitude-weighted equirectangular projection (ERP) coordinates. By leveraging relative depth to guide scale map upsampling and employing distortion-aware positional encoding, our approach achieves cross-camera generalization within a single model—without requiring multi-domain training or camera-specific architectures. The method sets new state-of-the-art results across multiple datasets, significantly outperforming existing approaches.