design lightweight depth models

Designs and implements neural architectures and model families for monocular depth estimation that are optimized to run in real time on mobile and edge hardware, producing variants across computation budgets. Analyzes and trades off accuracy, GFLOPS/latency and memory to meet deployment constraints while preserving robustness and zero‑shot generalization without relying on extra sensors.

designlightweightdepthmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of achieving efficient, low-power RGB-D affordance segmentation for wearable robots on embedded platforms. The authors propose two key innovations: first, a hardware-aware neural architecture search space specifically designed for depth-informed fusion; and second, a lightweight preprocessing layer that aligns depth maps with RGB inputs, enabling seamless integration with existing lightweight networks and a tailored fine-tuning strategy. As the first work to systematically explore RGB-D fusion for embedded affordance segmentation, this research achieves Pareto-optimal trade-offs between performance and power consumption on a Jetson Nano platform paired with a RealSense camera. The approach significantly outperforms current lightweight methods on real-world datasets and enables real-time operation under battery power.

affordance segmentationembedded deviceshardware constraints

This work addresses the challenge of balancing perception accuracy and computational efficiency in remote autonomous driving systems, which are constrained by limited computing resources, power budgets, and sensor capabilities on embedded platforms. The authors propose a context-adaptive monocular depth estimation method that, for the first time, closes the loop between perceptual fidelity and navigation task requirements. Their approach employs a slimmable neural network that dynamically adjusts computational complexity, activating high-fidelity inference only in critical scenarios. Evaluated on an open-source platform built with AirSim and Jetson Orin Nano, the system achieves a 16.1% reduction in power consumption, a 74.8% decrease in inference latency, and a 75.0% drop in overall energy usage compared to static baselines, while simultaneously improving navigation accuracy by 7.43% and reducing sensor data acquisition volume by 9.67%.

autonomous vehiclescomputational efficiencyembedded systems

Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation

Jul 15, 2025
ZX
Zhen Xu
🏛️ Zhejiang University | Shanghai AI Lab | Shenzhen University

Existing visual depth estimation methods suffer from poor generalization, low stability, and reliance on small-scale, domain-specific training data. Method: This paper proposes the “Depth Foundation Model” paradigm—a unified framework designed for strong zero-shot cross-scene transferability. It integrates diverse input modalities (monocular, stereo, multi-view, and video sequences) and employs self-supervised and weakly supervised learning strategies to train a high-capacity, modular neural architecture on large-scale heterogeneous datasets. Contribution/Results: We formally define the Depth Foundation Model concept and its technical roadmap for the first time; develop a scalable training framework and standardized evaluation benchmark; and demonstrate significant improvements in robustness and accuracy on unseen scenes. The resulting model enables high-resolution, environment-robust, and cost-effective depth perception—advancing applications in 3D reconstruction, autonomous driving, and AR/VR.

Developing cost-effective vision-based depth estimation methodsExploring large-scale datasets for robust depth foundation modelsImproving generalization and stability in depth estimation models

This work addresses the high computational cost of foundation models for monocular depth estimation, which hinders their real-time deployment on edge devices, and the underutilization of inter-frame redundancy in existing approaches. To overcome these limitations, we propose AsyncMDE, the first method to introduce an asynchronous spatial memory mechanism. It decouples computation by leveraging a large background model to generate high-quality spatial features while a lightweight foreground model performs real-time inference. Cross-frame feature reuse is enabled through an autoregressive memory update and a complementary fusion strategy. With only 3.83 million parameters, AsyncMDE achieves 237 FPS on an RTX 4090—recovering 77% of the foundation model’s accuracy—and 161 FPS on a Jetson AGX Orin, significantly outperforming current state-of-the-art methods.

computational redundancyedge deploymentfoundation model

This work addresses the vulnerability of lightweight monocular depth estimation models to domain shift in dynamic environments and their performance saturation due to static training paradigms. To overcome these limitations, the authors propose an online active learning framework featuring a closed-loop predict–evaluate–correct mechanism that actively selects high-informativeness samples from incoming visual streams for real-time model updating. By integrating selective plasticity with Elastic Weight Consolidation (EWC), the approach enables localized parameter adaptation while preserving globally learned knowledge, thereby breaking through the static optimization bottleneck inherent in compact architectures. Implemented on MobileNetV3-Small, the method achieves competitive accuracy with approximately 75% reduction in computational cost, demonstrating the critical role of controlled parameter plasticity in enabling effective dynamic adaptation.

domain shiftedge devicesmodel plasticity

Latest Papers

What's happening recently
View more

This work addresses the limited cross-domain generalization of existing lightweight monocular depth estimation models and the deployment challenges of high-accuracy foundation models on resource-constrained devices. To bridge this gap, we propose ZipDepth, which introduces large-scale multi-domain knowledge distillation into a compact architecture for the first time, coupled with an efficient reparameterizable encoder-decoder design. With only 6.1 million parameters, ZipDepth substantially narrows the accuracy gap with large models while achieving state-of-the-art zero-shot cross-domain performance across five benchmarks. The method strikes an optimal balance between generalization and inference efficiency, enabling real-time deployment ranging from server-grade GPUs to low-power edge devices.

domain shiftefficient deploymentlightweight models

This work addresses the limitation of local convolutions in monocular depth estimation, which struggle to capture long-range spatial dependencies. To overcome this, the authors propose GraphDepth, a novel architecture that integrates GraphSAGE graph neural networks into the multi-scale feature layers of a ResNet-101 U-Net to explicitly model global spatial relationships. Key innovations include a scalable batch-parallel graph construction strategy, multi-scale GNN integration, channel-attention-gated skip connections, and heteroscedastic uncertainty estimation. The method achieves near state-of-the-art accuracy on NYU Depth V2 at 25 FPS with only 3.8 GB GPU memory and sets a new best RMSE of 8.24 m on the WHU Aerial dataset, demonstrating exceptional cross-domain generalization capability.

computational efficiencycross-domain generalizationglobal context modeling

This work addresses the challenge of achieving efficient and low-power self-supervised monocular depth estimation on resource-constrained devices by proposing XiDepth, a novel network architecture built upon lightweight XiNet operator blocks. Departing from computationally expensive components such as depthwise separable convolutions and attention mechanisms, XiDepth leverages a self-supervised learning paradigm combined with an optimized feature extraction module to significantly reduce both parameter count and computational overhead while maintaining competitive accuracy. Experimental results demonstrate that XiDepth achieves state-of-the-art performance on the KITTI benchmark with only 0.8 million parameters. When deployed on a Raspberry Pi 4, it reduces FLOPs by 40% and energy consumption by 35% compared to existing approaches, offering high compatibility and deployment efficiency for edge platforms.

computational efficiencyembedded systemslightweight network

Hot Scholars

TL

Tianye Li

NVIDIA Research
Computer GraphicsComputer VisionDigital Humans
SZ

Siming Zheng

UCAS, vivo
AIGC,Low-level visionComputational photography,Snapshot Compressive Imaging,Deep Learning
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
DC

Deng Cai

Professor of Computer Science, Zhejiang University
Machine learningComputer visionData miningInformation retrieval
WM

Wenjun Mei

Peking University
multi-agent systemsnetwork dynamics and gamesmathematical sociology