🤖 AI Summary
This study addresses the high computational demands associated with classifying whole slide images (WSIs) of gastric adenocarcinoma by proposing a multimodal knowledge distillation framework. Without relying on Transformers or large language models, the method constructs a teacher model that integrates pretrained WSI encoders with clinical text features via low-rank fusion. The acquired multimodal knowledge is then distilled into a lightweight, image-only student model. Evaluated on the PatchGastric dataset, the proposed framework surpasses existing state-of-the-art methods by at least 3.35% in average accuracy. It achieves this superior performance while significantly reducing inference overhead, demonstrating an effective balance between high precision and computational efficiency. The source code has been made publicly available.
📝 Abstract
Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at https://github.com/helomelo1/MKD-LMF.