🤖 AI Summary
Intraoperative endoscopic videos suffer from physiological artifacts and noise, hindering robust cross-modal alignment with preoperative clean virtual images (e.g., CT/MRI 3D renderings). To address this, we propose a “local–global” disentangled image translation framework: a local module suppresses physiological artifacts and noise, while a global module enables robust style transfer from the real endoscopic domain to the virtual reference domain. We further design a noise-robust contrastive learning-based feature extractor to enhance pixel- and semantic-level cross-domain correspondence stability. Our method integrates domain adaptation, feature disentanglement modeling, and multi-source clinical data co-training. Evaluated on public and institutional datasets, our approach significantly outperforms existing state-of-the-art methods—achieving a 12.7% improvement in image matching accuracy under severe artifact conditions and reducing pose estimation mean error by 23.4%. This establishes a highly robust foundation for real-time endoscopic navigation via cross-modal alignment.
📝 Abstract
Domain adaptation, which bridges the distributions across different modalities, plays a crucial role in multimodal medical image analysis. In endoscopic imaging, combining pre-operative data with intra-operative imaging is important for surgical planning and navigation. However, existing domain adaptation methods are hampered by distribution shift caused by in vivo artifacts, necessitating robust techniques for aligning noisy and artifact abundant patient endoscopic videos with clean virtual images reconstructed from pre-operative tomographic data for pose estimation during intraoperative guidance. This paper presents an artifact-resilient image translation method and an associated benchmark for this purpose. The method incorporates a novel ``local-global'' translation framework and a noise-resilient feature extraction strategy. For the former, it decouples the image translation process into a local step for feature denoising, and a global step for global style transfer. For feature extraction, a new contrastive learning strategy is proposed, which can extract noise-resilient features for establishing robust correspondence across domains. Detailed validation on both public and in-house clinical datasets has been conducted, demonstrating significantly improved performance compared to the current state-of-the-art.