FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of multimodal alignment in federated biomedical settings, where data heterogeneity, label scarcity, and privacy constraints pose significant obstacles. To this end, we propose a lightweight federated few-shot classification framework that leverages Vision Mamba and Cross Mamba encoders to jointly optimize soft prompts and communication prompts. This approach achieves cross-modal feature alignment without relying on external large language models. By fine-tuning only the prompt parameters, the method substantially reduces computational overhead while effectively capturing cross-modal interaction structures. Extensive experiments demonstrate that the proposed framework consistently outperforms baseline methods across multiple biomedical datasets, achieving an average model size reduction of 1.96×. Overall, this work enables efficient and privacy-preserving multimodal learning for federated biomedical applications.
📝 Abstract
Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains. Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage. This is especially critical in biomedical applications under federated constraints, where data sharing is restricted, labeled data is scarce at each site, and it differs widely across institutions, leading to substantial statistical heterogeneity. To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy. With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance. Importantly, this eliminates the need to use an external Large Language Model (LLM) for feature alignment. The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed. The proposed framework delivers stable performance relative to the baselines and is, on average, 1.96 times lighter. The corresponding script will be made available soon.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Few-shot Classification
Vision-Language Models
Biomedical Imaging
Statistical Heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Prompt Learning
State Space Model
Few-shot Classification
Vision-Language Models
Lightweight
🔎 Similar Papers
A
Ankita Das
Department of Computer Science and Engineering, Indian Institute of Technology Hyderabad
A
Ambarish Parthasarathy
Department of Artificial Intelligence, Indian Institute of Technology Hyderabad
Sumohana S. Channappayya
Sumohana S. Channappayya
IIT Hyderabad
Image and Video Quality AssessmentMedical ImagingMachine Learning
C
C. Krishna Mohan
Department of Computer Science and Engineering, Indian Institute of Technology Hyderabad