🤖 AI Summary
This study addresses the poor generalization of learning-based depth estimation under novel UAV viewpoints and the geometric degeneracy inherent in conventional triangulation. To overcome these limitations, this work proposes a training-free framework based on epipolar transfer that leverages camera motion to convert temporally sequential monocular frames into virtual stereo pairs with variable baselines. By jointly modeling epipolar geometry and temporal correspondences for triangulation, the proposed approach effectively circumvents degenerate configurations. Notably, this method requires no training data yet achieves accuracy comparable to direct triangulation while significantly outperforming mainstream deep learning baselines across both indoor and outdoor scenes. Consequently, it substantially enhances the robustness of depth estimation from UAV perspectives, offering a practical and reliable solution for aerial 3D perception tasks.
📝 Abstract
Reliable 3D spatial understanding is essential for autonomous navigation, obstacle avoidance, and scene reconstruction. While state-of-the-art learned depth estimation techniques achieve high accuracy in-distribution, they often generalize poorly to novel viewpoints and altitudes. This paper presents a geometrically derived, training-free depth estimation method using epipolar transfer with only two monocular images and camera pose estimates. By leveraging camera motion to synthesize a virtual stereo pair with a freely chosen baseline, our approach transforms temporal correspondence into a stereo triangulation task while mitigating geometric degeneracies inherent to direct two-view triangulation. Validated across outdoor drone flights (to a maximum range of approximately 90\,m) and indoor OptiTrack environments against LiDAR ground truth, the method achieves an indoor AbsRel of 0.092 and $\delta<1.25$ of 0.940, comparable to direct triangulation (AbsRel 0.073) while retaining valid depth over a larger fraction of challenging scenes, and substantially outperforms off-the-shelf learning-based baselines such as ZoeDepth (AbsRel 0.225) and Depth Anything V2 (AbsRel 0.570), which are not trained or fine-tuned for this domain, with no training data required.