CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CMambaDepth框架,通过双向通道曼巴和混合注意力模块解决单目深度估计中的跨尺度信息交互低效问题及局部与全局空间建模平衡难题。
📝 Abstract
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and global spatial modeling. In this paper, we propose CMambaDepth, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation. Specifically, Bidirectional Channel Mamba (Bi-CMamba) aligns encoder features across scales and enables bidirectional information exchange among ordered scale groups. Unidirectional Channel Mamba (Uni-CMamba) progressively aggregates decoder features and retains fine-grained scale groups through a group selection mechanism for subsequent fusion. Furthermore, a Hybrid Attention Module (HAM) is introduced to combine large-kernel local context and Manhattan self-attention for complementary spatial modeling. Experimental results demonstrate that our method achieves highly competitive performance. Specifically, our model achieves an AbsRel of 0.094 and an RMSE of 4.156 on KITTI, and an AbsRel of 0.140 on DDAD. In the zero-shot cross-dataset generalization test on NYUv2, it attains an AbsRel of 0.232, outperforming the baseline RA-Depth by 7.2%.
Problem

Research questions and friction points this paper is trying to address.

monocular depth estimation
cross-scale information interaction
spatial modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised
Channel Mamba
Hybrid Attention
Multi-scale feature fusion
Fine-grained contextual modeling
💼 Related Jobs
No related jobs found.
Xuezhi Xiang
Xuezhi Xiang
School of Information and Communication Engineering, Harbin Engineering University, Harbin 150001, China
J
Jiayao Liu
School of Information and Communication Engineering, Harbin Engineering University, Harbin 150001, China
H
Heqi Xiang
Department of Computer Science, University of Toronto, Toronto ON M5S 2E4, Canada
Yuqi Hu
Yuqi Hu
HKUST(GZ)
Machine LearningCV&CGGenerative AI
Y
Yiming Chen
School of Information and Communication Engineering, Harbin Engineering University, Harbin 150001, China
S
Shanjun Zhang
Department of Computer Science, Kanagawa University, Kanagawa 221-8686, Japan