FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决病理图像标注成本高和少样本学习问题,提出FFM-CP框架,通过融合多个预训练模型并采用正交普罗克鲁斯特变换对齐特征,提高分类性能。
📝 Abstract
Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. Combining complementary pretrained representations is a potential approach to these limitations, yet learning an effective fusion from few labeled examples remains challenging. We introduce Few-shot Fusion Foundation Models of Computational Pathology (FFM-CP), which is a framework that combines multiple pathology vision-language models in the few-shot learning setting. The framework first aligns heterogeneous representations using a closed-form Orthogonal Procrustes transformation estimated from corresponding support images. This alignment preserves within-model feature geometry without training an additional alignment network. Within the aligned space, a unified graph enables information exchange across backbones by jointly refining support-image features and visual and textual class prototypes. These refined representations support complementary text-prototype and case-retrieval branches that capture semantic class knowledge and within-class visual variation, respectively. Each branch learns to combine predictions from all ordered backbone pairs, allowing queries encoded by one model to draw on evidence represented by another. We evaluate three backbone combinations on six histopathology datasets at 4, 8, and 16 shots per class. FFM-CP achieves higher mean macro-F1 than the strongest individually adapted member of each fused set in 50 of 54 comparisons. These findings suggest that combining complementary pretrained representations can improve histopathological classification when annotations are limited.
Problem

Research questions and friction points this paper is trying to address.

pathology vision-language models
few-shot learning
complementary pretrained representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-shot Learning
Cross-Backbone Fusion
Orthogonal Procrustes Transformation
Graph-based Information Exchange
Prototype Refinement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Anh-Tien Nguyen
Institute for Predictive Deep Learning for Medicine and Healthcare, Giessen University, Germany
T
Trung DQ. Dang
Department of Applied Mathematics and Computer Science, Technical University of Denmark, Denmark
N
Nghiem Tuong Diep
Department of machine learning, Mohamed bin Zayed University of Artificial Intelligence, UAE
B
Bui Ngoc Han Nguyen
Department of machine learning, Mohamed bin Zayed University of Artificial Intelligence, UAE
T
Tan-Ha Mai
Department of Computer Science and Information Engineering, National Taiwan University, Taiwan
M
Miriam Cindy Maurer
Department of Medical Informatics, University Medicine Gottingen, Germany
P
Phuong Hoa Nguyen
Faculty of Basic Medicine and Pharmacy, VNU University of Medicine and Pharmacy, Vietnam
T
Thi Thuy Uyen Nguyen
Department of Histology, Embryology, Pathology and Forensic Medicine, University of Medicine and Pharmacy, Hue University, Vietnam
Youngjun Park
Youngjun Park
Max Planck Institute for Biology of Ageing
Computational BiologyMachine learningBioinformatics
Daniel Sonntag
Daniel Sonntag
DFKI and University of Oldenburg
Interactive Machine LearningIntelligent User InterfacesMultimodal Interaction
D
Duy Minh Ho Nguyen
Max Planck Research School for Intelligent Systems (IMPRS-IS), Germany
Anne-Christin Hauschild
Anne-Christin Hauschild
University Professor at Justus-Liebig University Gießen
Machine LearningExplainable AIBioinformaticsBiomedical Data ScienceBiostatistics