Score
Designs and implements neural architectures and model families for monocular depth estimation that are optimized to run in real time on mobile and edge hardware, producing variants across computation budgets. Analyzes and trades off accuracy, GFLOPS/latency and memory to meet deployment constraints while preserving robustness and zero‑shot generalization without relying on extra sensors.
Monocular depth estimation (MDE) suffers from inconsistent scale due to the absence of metric ground truth, severely limiting its applicability in downstream tasks such as visual localization and 3D reconstruction. This work addresses zero-shot monocular metric depth estimation (MMDE), aiming to improve cross-scene generalization and boundary detail fidelity. We first systematically survey the evolution of scale-invariant approaches and establish zero-shot MMDE as the canonical benchmark paradigm. Our method innovatively integrates label-free data augmentation, image patch-based modeling, geometry-aware architectural optimization, and generative prior learning. Experiments demonstrate substantial improvements in depth scale consistency and structural integrity across multiple benchmarks. Under the zero-shot setting, our approach achieves superior calibration accuracy and generalization performance compared to prior methods. Furthermore, we identify key open challenges and promising directions for future research in metric monocular depth estimation.
This study addresses the challenge of achieving efficient, low-power RGB-D affordance segmentation for wearable robots on embedded platforms. The authors propose two key innovations: first, a hardware-aware neural architecture search space specifically designed for depth-informed fusion; and second, a lightweight preprocessing layer that aligns depth maps with RGB inputs, enabling seamless integration with existing lightweight networks and a tailored fine-tuning strategy. As the first work to systematically explore RGB-D fusion for embedded affordance segmentation, this research achieves Pareto-optimal trade-offs between performance and power consumption on a Jetson Nano platform paired with a RealSense camera. The approach significantly outperforms current lightweight methods on real-world datasets and enables real-time operation under battery power.
This work addresses the challenge of balancing perception accuracy and computational efficiency in remote autonomous driving systems, which are constrained by limited computing resources, power budgets, and sensor capabilities on embedded platforms. The authors propose a context-adaptive monocular depth estimation method that, for the first time, closes the loop between perceptual fidelity and navigation task requirements. Their approach employs a slimmable neural network that dynamically adjusts computational complexity, activating high-fidelity inference only in critical scenarios. Evaluated on an open-source platform built with AirSim and Jetson Orin Nano, the system achieves a 16.1% reduction in power consumption, a 74.8% decrease in inference latency, and a 75.0% drop in overall energy usage compared to static baselines, while simultaneously improving navigation accuracy by 7.43% and reducing sensor data acquisition volume by 9.67%.
Existing visual depth estimation methods suffer from poor generalization, low stability, and reliance on small-scale, domain-specific training data. Method: This paper proposes the “Depth Foundation Model” paradigm—a unified framework designed for strong zero-shot cross-scene transferability. It integrates diverse input modalities (monocular, stereo, multi-view, and video sequences) and employs self-supervised and weakly supervised learning strategies to train a high-capacity, modular neural architecture on large-scale heterogeneous datasets. Contribution/Results: We formally define the Depth Foundation Model concept and its technical roadmap for the first time; develop a scalable training framework and standardized evaluation benchmark; and demonstrate significant improvements in robustness and accuracy on unseen scenes. The resulting model enables high-resolution, environment-robust, and cost-effective depth perception—advancing applications in 3D reconstruction, autonomous driving, and AR/VR.
This work addresses the high computational cost of foundation models for monocular depth estimation, which hinders their real-time deployment on edge devices, and the underutilization of inter-frame redundancy in existing approaches. To overcome these limitations, we propose AsyncMDE, the first method to introduce an asynchronous spatial memory mechanism. It decouples computation by leveraging a large background model to generate high-quality spatial features while a lightweight foreground model performs real-time inference. Cross-frame feature reuse is enabled through an autoregressive memory update and a complementary fusion strategy. With only 3.83 million parameters, AsyncMDE achieves 237 FPS on an RTX 4090—recovering 77% of the foundation model’s accuracy—and 161 FPS on a Jetson AGX Orin, significantly outperforming current state-of-the-art methods.
This work addresses the vulnerability of lightweight monocular depth estimation models to domain shift in dynamic environments and their performance saturation due to static training paradigms. To overcome these limitations, the authors propose an online active learning framework featuring a closed-loop predict–evaluate–correct mechanism that actively selects high-informativeness samples from incoming visual streams for real-time model updating. By integrating selective plasticity with Elastic Weight Consolidation (EWC), the approach enables localized parameter adaptation while preserving globally learned knowledge, thereby breaking through the static optimization bottleneck inherent in compact architectures. Implemented on MobileNetV3-Small, the method achieves competitive accuracy with approximately 75% reduction in computational cost, demonstrating the critical role of controlled parameter plasticity in enabling effective dynamic adaptation.
本文综述了单目深度估计的发展,从早期基于学习的方法到基础模型的出现,探讨了相对和度量深度估计的关键挑战及解决方案。
This work addresses the limited cross-domain generalization of existing lightweight monocular depth estimation models and the deployment challenges of high-accuracy foundation models on resource-constrained devices. To bridge this gap, we propose ZipDepth, which introduces large-scale multi-domain knowledge distillation into a compact architecture for the first time, coupled with an efficient reparameterizable encoder-decoder design. With only 6.1 million parameters, ZipDepth substantially narrows the accuracy gap with large models while achieving state-of-the-art zero-shot cross-domain performance across five benchmarks. The method strikes an optimal balance between generalization and inference efficiency, enabling real-time deployment ranging from server-grade GPUs to low-power edge devices.
This work addresses the limitation of local convolutions in monocular depth estimation, which struggle to capture long-range spatial dependencies. To overcome this, the authors propose GraphDepth, a novel architecture that integrates GraphSAGE graph neural networks into the multi-scale feature layers of a ResNet-101 U-Net to explicitly model global spatial relationships. Key innovations include a scalable batch-parallel graph construction strategy, multi-scale GNN integration, channel-attention-gated skip connections, and heteroscedastic uncertainty estimation. The method achieves near state-of-the-art accuracy on NYU Depth V2 at 25 FPS with only 3.8 GB GPU memory and sets a new best RMSE of 8.24 m on the WHU Aerial dataset, demonstrating exceptional cross-domain generalization capability.
This work addresses the challenge of achieving efficient and low-power self-supervised monocular depth estimation on resource-constrained devices by proposing XiDepth, a novel network architecture built upon lightweight XiNet operator blocks. Departing from computationally expensive components such as depthwise separable convolutions and attention mechanisms, XiDepth leverages a self-supervised learning paradigm combined with an optimized feature extraction module to significantly reduce both parameter count and computational overhead while maintaining competitive accuracy. Experimental results demonstrate that XiDepth achieves state-of-the-art performance on the KITTI benchmark with only 0.8 million parameters. When deployed on a Raspberry Pi 4, it reduces FLOPs by 40% and energy consumption by 35% compared to existing approaches, offering high compatibility and deployment efficiency for edge platforms.
MVP通过预测未来运动矢量并在运动域中推测视觉结果,减少连续视觉系统中的端到端延迟和能耗。