FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inherent lack of absolute scale in monocular depth estimation and the difficulty existing completion networks face in balancing geometric fidelity with inference efficiency. To this end, this work proposes a training-free, modular framework that leverages sparse LiDAR anchors to align a frozen foundation model (MoGe-2), employing a structure-preserving solver to generate dense metric depth maps. The proposed approach validates anchor consistency and preserves fine-grained geometric structures without requiring any fine-tuning. Experimental results demonstrate that the method achieves a 24% reduction in out-of-domain error, a 92% decrease in normal noise, and a 15-fold improvement in inference speed. These findings indicate a synergistic advancement in accuracy, generalization capability, and computational efficiency for metric depth completion.
πŸ“ Abstract
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-completion networks recover metric depth at the cost of geometric fidelity, cross-domain robustness, or speed. We present FounRef, a training-free method that aligns a frozen monocular foundation prior with sparse metric anchors to produce dense metric depth. FounRef is modular by design: its depth prior, anchor source, and refinement solver can each be replaced independently. We instantiate FounRef with MoGe-2 and LiDAR anchors. FounRef validates each anchor against the prior's dense depth prediction, rejecting inconsistencies caused by cross-sensor misalignment that geometry-only filters cannot detect. It then applies global and local metric corrections through a structure-preserving solver, retaining the prior's fine-grained geometry. FounRef requires no task-specific training and operates out of the box across unfamiliar cameras and scenes. On out-of-domain data, it delivers up to 24% lower depth error, 92% lower surface-normal noise, and almost 15x faster inference than DMD3C, a state-of-the-art depth-completion network. By decoupling metric alignment from geometry prediction, FounRef provides an accurate, geometrically faithful, and efficient approach to dense metric depth that can directly benefit from future advances in foundation models and metric sensors.
Problem

Research questions and friction points this paper is trying to address.

dense metric depth
monocular foundation models
depth completion
geometric fidelity
cross-domain robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free
Metric Depth Refinement
Structure-Preserving Solver
Monocular Foundation Models
Modular Design
πŸ’Ό Related Jobs
No related jobs found.