FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification
This study addresses the challenge of multimodal alignment in federated biomedical settings, where data heterogeneity, label scarcity, and privacy constraints pose significant obstacles. To this end, we propose a lightweight federated few-shot classification framework that leverages Vision Mamba and Cross Mamba encoders to jointly optimize soft prompts and communication prompts. This approach achieves cross-modal feature alignment without relying on external large language models. By fine-tuning only the prompt parameters, the method substantially reduces computational overhead while effectively capturing cross-modal interaction structures. Extensive experiments demonstrate that the proposed framework consistently outperforms baseline methods across multiple biomedical datasets, achieving an average model size reduction of 1.96×. Overall, this work enables efficient and privacy-preserving multimodal learning for federated biomedical applications.