A Hybrid Mamba-SAM Architecture for Efficient 3D Medical Image Segmentation

📅 2026-01-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance limitations of general-purpose foundation models like SAM in 3D medical image segmentation, which stem from domain shift, inherent 2D architectural constraints, and the high cost of full fine-tuning. To overcome these challenges, the authors propose an efficient dual-branch fusion architecture that freezes the SAM Vision Transformer (ViT) encoder to preserve its general-purpose representations while integrating a Mamba-based state space model for effective 3D modeling. Key innovations include Tri-Plane Mamba for implicit 3D representation, multi-frequency gated convolutions to enrich feature representation, and a lightweight adapter module enabling fast inference. Evaluated on the ACDC dataset, the dual-branch model achieves a Dice score of 0.906 (0.910 for myocardium, 0.971 for left ventricle), while the adapter variant attains 0.880 Dice at 4.77 FPS, demonstrating an excellent balance between accuracy and efficiency.

Technology Category

Computer Vision: 3D Computer VisionMachine Learning: Deep Neural Architectures and Foundation ModelsSearch and Optimization: Sampling/Simulation-based Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
Accurate segmentation of 3D medical images such as MRI and CT is essential for clinical diagnosis and treatment planning. Foundation models like the Segment Anything Model (SAM) provide powerful general-purpose representations but struggle in medical imaging due to domain shift, their inherently 2D design, and the high computational cost of fine-tuning. To address these challenges, we propose Mamba-SAM, a novel and efficient hybrid architecture that combines a frozen SAM encoder with the linear-time efficiency and long-range modeling capabilities of Mamba-based State Space Models (SSMs). We investigate two parameter-efficient adaptation strategies. The first is a dual-branch architecture that explicitly fuses general features from a frozen SAM encoder with domain-specific representations learned by a trainable VMamba encoder using cross-attention. The second is an adapter-based approach that injects lightweight, 3D-aware Tri-Plane Mamba (TPMamba) modules into the frozen SAM ViT encoder to implicitly model volumetric context. Within this framework, we introduce Multi-Frequency Gated Convolution (MFGC), which enhances feature representation by jointly analyzing spatial and frequency-domain information via 3D discrete cosine transforms and adaptive gating. Extensive experiments on the ACDC cardiac MRI dataset demonstrate the effectiveness of the proposed methods. The dual-branch Mamba-SAM-Base model achieves a mean Dice score of 0.906, comparable to UNet++ (0.907), while outperforming all baselines on Myocardium (0.910) and Left Ventricle (0.971) segmentation. The adapter-based TP MFGC variant offers superior inference speed (4.77 FPS) with strong accuracy (0.880 Dice). These results show that hybridizing foundation models with efficient SSM-based architectures provides a practical and effective solution for 3D medical image segmentation.
Problem

Research questions and friction points this paper is trying to address.

3D medical image segmentation
domain shift
computational cost
foundation models
Segment Anything Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mamba-SAM
State Space Models
3D Medical Image Segmentation
Parameter-Efficient Adaptation
Multi-Frequency Gated Convolution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mohammadreza Gholipour Shahraki
Department of Electrical and Computer Engineering, Isfahan University of Technology, Isfahan, Iran
Mehdi Rezaeian
Mehdi Rezaeian
Associate Professor of Computer Engineering Department, Yazd University
Computer visionDeep learningPhotogrammetry
M
Mohammad Ghasemzadeh
Department of Computer Engineering, Yazd University, Yazd, Iran