From Real Artifacts to Virtual Reference: A Robust Framework for Translating Endoscopic Images

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Intraoperative endoscopic videos suffer from physiological artifacts and noise, hindering robust cross-modal alignment with preoperative clean virtual images (e.g., CT/MRI 3D renderings). To address this, we propose a “local–global” disentangled image translation framework: a local module suppresses physiological artifacts and noise, while a global module enables robust style transfer from the real endoscopic domain to the virtual reference domain. We further design a noise-robust contrastive learning-based feature extractor to enhance pixel- and semantic-level cross-domain correspondence stability. Our method integrates domain adaptation, feature disentanglement modeling, and multi-source clinical data co-training. Evaluated on public and institutional datasets, our approach significantly outperforms existing state-of-the-art methods—achieving a 12.7% improvement in image matching accuracy under severe artifact conditions and reducing pose estimation mean error by 23.4%. This establishes a highly robust foundation for real-time endoscopic navigation via cross-modal alignment.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningIntelligent Robots: Localization, Mapping, and Navigation

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Domain adaptation, which bridges the distributions across different modalities, plays a crucial role in multimodal medical image analysis. In endoscopic imaging, combining pre-operative data with intra-operative imaging is important for surgical planning and navigation. However, existing domain adaptation methods are hampered by distribution shift caused by in vivo artifacts, necessitating robust techniques for aligning noisy and artifact abundant patient endoscopic videos with clean virtual images reconstructed from pre-operative tomographic data for pose estimation during intraoperative guidance. This paper presents an artifact-resilient image translation method and an associated benchmark for this purpose. The method incorporates a novel ``local-global'' translation framework and a noise-resilient feature extraction strategy. For the former, it decouples the image translation process into a local step for feature denoising, and a global step for global style transfer. For feature extraction, a new contrastive learning strategy is proposed, which can extract noise-resilient features for establishing robust correspondence across domains. Detailed validation on both public and in-house clinical datasets has been conducted, demonstrating significantly improved performance compared to the current state-of-the-art.
Problem

Research questions and friction points this paper is trying to address.

Aligning noisy endoscopic videos with clean virtual images
Addressing distribution shift from in vivo artifacts
Improving pose estimation for intraoperative guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Local-global translation framework for denoising
Noise-resilient feature extraction via contrastive learning
Artifact-resilient image translation method
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shanghai Jiao Tong University | Shanghai Chest Hospital
J
Junyang Wu
Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, CHINA.
F
Fangfang Xie
Shanghai Chest Hospital, Shanghai, CHINA
J
Jiayuan Sun
Shanghai Chest Hospital, Shanghai, CHINA
Yun Gu
Yun Gu
Shanghai Jiao Tong University
Medical Image AnalysisComputer-Assisted Intervention
G
Guang-Zhong Yang
Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, CHINA.