Mixture-of-Shape-Experts (MoSE): End-to-End Shape Dictionary Framework to Prompt SAM for Generalizable Medical Segmentation

📅 2025-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Medical image single-domain generalization (SDG) suffers from weak shape prior modeling and limitations of dictionary-based approaches—including susceptibility to overfitting and poor compatibility with large foundation models. To address these issues, we propose an end-to-end trainable Mixture-of-Experts (MoE) shape dictionary framework: it decomposes shape priors into multiple specialized shape experts and dynamically fuses them via sparse gating to generate structured shape maps as prompts for Segment Anything Model (SAM). This work is the first to integrate the MoE mechanism into shape dictionary learning, enabling bidirectional collaboration and joint optimization between shape experts and SAM. Our method simultaneously alleviates offline dictionary capacity constraints and overfitting, achieving significant improvements in generalization performance across multi-center, multi-device, and multi-protocol medical datasets—outperforming existing SDG methods while maintaining low parameter count and high inference efficiency.

Technology Category

Machine Learning: Mixture of Experts (MoE)Computer Vision: Multi-modal VisionSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSystems and Infrastructure for Web, Mobile and WoT: Web applications in cross-disciplinary domains and verticals such as mixed reality, smart cities, and digital health
📝 Abstract
Single domain generalization (SDG) has recently attracted growing attention in medical image segmentation. One promising strategy for SDG is to leverage consistent semantic shape priors across different imaging protocols, scanner vendors, and clinical sites. However, existing dictionary learning methods that encode shape priors often suffer from limited representational power with a small set of offline computed shape elements, or overfitting when the dictionary size grows. Moreover, they are not readily compatible with large foundation models such as the Segment Anything Model (SAM). In this paper, we propose a novel Mixture-of-Shape-Experts (MoSE) framework that seamlessly integrates the idea of mixture-of-experts (MoE) training into dictionary learning to efficiently capture diverse and robust shape priors. Our method conceptualizes each dictionary atom as a shape expert, which specializes in encoding distinct semantic shape information. A gating network dynamically fuses these shape experts into a robust shape map, with sparse activation guided by SAM encoding to prevent overfitting. We further provide this shape map as a prompt to SAM, utilizing the powerful generalization capability of SAM through bidirectional integration. All modules, including the shape dictionary, are trained in an end-to-end manner. Extensive experiments on multiple public datasets demonstrate its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Enhancing single domain generalization in medical image segmentation
Overcoming limitations of shape prior encoding in dictionary learning
Integrating shape priors with SAM for robust segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Shape-Experts integrates MoE into dictionary learning
Gating network dynamically fuses shape experts sparsely
Shape map prompts SAM for bidirectional generalization
🔎 Similar Papers
No similar papers found.
J
Jia Wei
Dept. of Radiology and Biomedical Imaging, Yale University, New Haven, USA; Dept. of Biomedical Informatics and Data Science, Yale University, New Haven, USA
Xiaoqi Zhao
Xiaoqi Zhao
Yale University
Computer VisionAi4IndustryAi4HealthEV Safety
Jonghye Woo
Jonghye Woo
Associate Professor of Radiology, Harvard Medical School | MGH
Medical Image AnalysisMedical ImagingComputer VisionMachine LearningSpeech
J
J. Ouyang
Dept. of Radiology and Biomedical Imaging, Yale University, New Haven, USA; Dept. of Biomedical Informatics and Data Science, Yale University, New Haven, USA
G
G. Fakhri
Dept. of Radiology and Biomedical Imaging, Yale University, New Haven, USA; Dept. of Biomedical Informatics and Data Science, Yale University, New Haven, USA
Qingyu Chen
Qingyu Chen
Biomedical Informatics & Data Science, Yale University; NCBI-NLM, National Institutes of Health
Text miningMachine learningData curationBioNLPMedical Imaging Analysis
X
Xiaofeng Liu
Dept. of Radiology and Biomedical Imaging, Yale University, New Haven, USA; Dept. of Biomedical Informatics and Data Science, Yale University, New Haven, USA