EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of geometric details in high-resolution images and the substantial inference overhead caused by existing VAE-based reconstruction in monocular depth estimation. To overcome these limitations, this work proposes a method that integrates latent diffusion priors with pixel-level generation. The core innovation lies in a two-stage PiD decoder architecture that bypasses the original VAE decoder by employing a latent branch to process low-resolution inputs and a pixel branch to directly generate high-resolution outputs, coupled with a multi-resolution progressive training strategy. Extensive evaluations demonstrate that the proposed approach achieves state-of-the-art performance across five standard benchmarks and the Synth4K dataset. Furthermore, it significantly accelerates inference while effectively preserving fine-grained structures and object boundaries.
📝 Abstract
Recovering detailed geometry from high-resolution images is critical for precise perception of the surroundings and objects. However, existing methods which use latent-space modeling and VAE reconstruction can compromise geometric details. Furthermore, decoding from latent codes introduces substantial inference overhead. To address those issues, we present EagleDepth, an efficient framework for high-resolution monocular depth estimation that combines the geometric priors of latent diffusion with fine-grained pixel-space generation. Our key idea is to retain depth-aware latent representations as guidance while generating the final depth map directly in pixel space. We train the latent and pixel components sequentially: first, we fine-tune a pretrained latent diffusion model using paired RGB--depth supervision; then, we adapt a pretrained pixel diffusion decoder, PiD, to predict depth conditioned on the learned features. Training of the pixel component starts at 1024 resolution and continues across multiple resolutions up to 4K. The latent branch processes resized, lower-resolution RGB images, while the pixel branch generates depth at the target resolution, bypassing the original VAE decoder. This design preserves learned geometric knowledge without requiring the latent backbone to operate at the output resolution. On five commonly used depth estimation datasets and the high-resolution Synth4K dataset, our framework achieves state-of-the-art depth estimation performance, with faster inference and better preservation of fine structures and object boundaries.
Problem

Research questions and friction points this paper is trying to address.

monocular depth estimation
high-resolution images
fine-grained geometry
inference overhead
latent diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Monocular Depth Estimation
Pixel Diffusion Decoder
Latent Diffusion Model
High-Resolution
Fine-Grained Geometry
💼 Related Jobs
No related jobs found.
B
Bowen Chai
Shanghai Jiao Tong University
T
Tianbao Zhang
Shanghai Jiao Tong University
S
Shuyu Wu
Shanghai Jiao Tong University
D
Dexin Zuo
Shanghai Jiao Tong University
Z
Zhaoxin Fan
Beihang University
Danping Zou
Danping Zou
Professor, Shanghai Jiao Tong University
Visual SLAMRobotic VisionVision-based navigation