🤖 AI Summary
This study addresses the high cost and time-consuming nature of spatial gene expression profiling, noting that existing methods predominantly rely on image features while neglecting the rich semantic information embedded in gene textual descriptions. To bridge this gap, this work pioneers the integration of gene textual annotations into spatial transcriptomics prediction by proposing a multimodal fusion framework based on cross-attention mechanisms. Through deep interaction between a text encoder and visual features, the method achieves precise gene expression prediction. Experimental results demonstrate that the proposed approach significantly outperforms multiple baseline models, effectively overcoming the limitations inherent in relying solely on visual representations. Ultimately, this research establishes a novel paradigm for cost-effective and efficient spatial transcriptomics analysis.
📝 Abstract
Spatial transcriptomics enables spatially resolved gene expression analysis from slide-level images while preserving morphological features, providing valuable information for studying disease mechanisms and developing treatments. However, spatial gene expression profiling typically requires expensive and time-consuming tests. While existing image-based prediction optimizations mostly revolve around including positional embeddings and further image-based changes, text-based optimizations remain relatively unexplored. We present GATE-ST, which incorporates text-based inputs into image-based spatial gene expression predictions. With this approach, generated text descriptions of genes are utilized to better spatial transcriptomics prediction results. Gene summaries are put through a text encoder, generating embeddings that integrate with image embeddings through cross-attention layers to align with morphological features. We demonstrate the effectiveness of such text inputs by benchmarking performance against random gene embeddings and multiple other image-text fusion architectures, and show that GATE-ST outperforms these alternatives. Our results demonstrate the effectiveness of GATE-ST in pathology imaging, which may greatly reduce the time and cost of accurate spatial transcriptomic predictions, proving the potential of text-guided spatial gene expression prediction.