🤖 AI Summary
This work addresses the challenge of generalizing depth and ego-motion estimation in endoscopic scenes, where illumination artifacts and feature diversity hinder performance. To tackle this, the authors propose EndoMINI, a self-supervised framework that introduces Mixture of Low-Rank Experts (MiLoRE) for parameter-efficient fine-tuning. EndoMINI further integrates an intrinsic image decomposition network with an alignment loss to effectively suppress specular reflections and enhance cross-scene adaptability. Evaluated on the SCARED, Hamlyn, and SERV-CT datasets, EndoMINI consistently achieves superior depth estimation accuracy compared to state-of-the-art methods, demonstrating significantly improved generalization and robustness.
📝 Abstract
Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various endoscopic scenes are still challenges for generalizable depth estimation and ego-motion estimation. Based on this, a novel self-supervised framework, EndoMINI, is proposed for depth estimation in endoscopic scenes. Specifically, mixture of low-rank experts (MiLoRE) is proposed to perform parameter-efficient fine-tuning, which can also boost the model adaptation to scenes with different characteristics. Meanwhile, an intrinsic image alignment (IIA) is introduced into the training loss to alleviate the influence of light reflectance in endoscopy with a novel intrinsic image decomposition network. The proposed method is evaluated on SCARED datasets for supervised depth estimation, and two endoscopic datasets, Hamlyn and SERV-CT, for zero-shot depth estimation, compared with state-of-the-art works as well. The experimental results demonstrate outstanding performance of the proposed model and the effects of the main contributions.