🤖 AI Summary
This study addresses the reliance on costly RNA sequencing for identifying immunotherapy biomarkers in gastric cancer by proposing CCMIL, a cross-modal retrieval framework. This method integrates cross-modal contrastive multi-instance learning, supervised contrastive objectives, and latent space alignment to directly infer molecular features from routine H&E-stained histopathology slides, eliminating the need for genomic sequencing during inference. Furthermore, this work constructs an interpretable case-level retrieval engine that effectively captures the continuous phenotypic spectrum of tumor inflammation and generates clinically interpretable attention heatmaps. By achieving competitive performance in downstream classification tasks, the proposed framework establishes a novel paradigm for cost-effective precision medicine.
📝 Abstract
Gastric Adenocarcinoma is a leading cause of cancer mortality. Although"Inflamed/Non-Inflamed"subtypes have been proposed to predict immunotherapy response, their identification relies on a costly 10-gene RNA signature. We propose a Cross-modal Contrastive Multiple Instance Learning (CCMIL) framework for cross-modal retrieval, imputing these molecular signatures directly from standard Hematoxylin&Eosin (H&E) slides. By leveraging a supervised contrastive objective, CCMIL aligns visual morphological patterns with molecular phenotypes into a shared latent space. This establishes an interpretable search-by-case retrieval engine, enabling pathologists to query a whole slide image to surface transcriptomically coherent neighbors and approximate RNA signatures without genomic sequencing at inference. Our results demonstrate that this retrieval-first approach captures the continuous phenotypic spectrum of tumor inflammation and yields clinically interpretable attention heatmaps. Furthermore, the learned representation also supports competitive downstream classification, providing a practical molecular pre-screening strategy.