EpiTransfer: Sparse, Training-Free Long-Range Depth Estimation from Temporal Monocular Aerial Frames

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the poor generalization of learning-based depth estimation under novel UAV viewpoints and the geometric degeneracy inherent in conventional triangulation. To overcome these limitations, this work proposes a training-free framework based on epipolar transfer that leverages camera motion to convert temporally sequential monocular frames into virtual stereo pairs with variable baselines. By jointly modeling epipolar geometry and temporal correspondences for triangulation, the proposed approach effectively circumvents degenerate configurations. Notably, this method requires no training data yet achieves accuracy comparable to direct triangulation while significantly outperforming mainstream deep learning baselines across both indoor and outdoor scenes. Consequently, it substantially enhances the robustness of depth estimation from UAV perspectives, offering a practical and reliable solution for aerial 3D perception tasks.
📝 Abstract
Reliable 3D spatial understanding is essential for autonomous navigation, obstacle avoidance, and scene reconstruction. While state-of-the-art learned depth estimation techniques achieve high accuracy in-distribution, they often generalize poorly to novel viewpoints and altitudes. This paper presents a geometrically derived, training-free depth estimation method using epipolar transfer with only two monocular images and camera pose estimates. By leveraging camera motion to synthesize a virtual stereo pair with a freely chosen baseline, our approach transforms temporal correspondence into a stereo triangulation task while mitigating geometric degeneracies inherent to direct two-view triangulation. Validated across outdoor drone flights (to a maximum range of approximately 90\,m) and indoor OptiTrack environments against LiDAR ground truth, the method achieves an indoor AbsRel of 0.092 and $\delta<1.25$ of 0.940, comparable to direct triangulation (AbsRel 0.073) while retaining valid depth over a larger fraction of challenging scenes, and substantially outperforms off-the-shelf learning-based baselines such as ZoeDepth (AbsRel 0.225) and Depth Anything V2 (AbsRel 0.570), which are not trained or fine-tuned for this domain, with no training data required.
Problem

Research questions and friction points this paper is trying to address.

depth estimation
monocular aerial frames
long-range
training-free
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free depth estimation
epipolar transfer
virtual stereo pair
long-range monocular depth
geometric degeneracy mitigation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Diksha Aggarwal
Department of Mechanical Engineering, Virginia Tech, Blacksburg, Virginia, USA.
R
Rutvik Dagadkhair
Department of Mechanical Engineering, Virginia Tech, Blacksburg, Virginia, USA.
Sanjana Srivastava
Sanjana Srivastava
Computer Science Ph.D. student, Stanford University
Cognitive scienceEmbodied AIComputer vision
B
Bradley Denby
Department of Aerospace and Ocean Engineering, Virginia Tech, Blacksburg, Virginia.
Kevin Kochersberger
Kevin Kochersberger
Department of Mechanical Engineering, Virginia Tech, Blacksburg, Virginia, USA.