Score
Combining learned cell embeddings and annotations with spatial transcriptomics data to validate biological signals and detect spatially localized activity or depot-specific differences, enabling detection and biological validation of signaling localization.
This study addresses the challenge of unified modeling of single-cell multimodal data—encompassing morphology, gene expression, and spatial coordinates—by proposing the first cross-scale, spatially aware generative framework. Methodologically, it introduces a multi-scale Transformer architecture that ingests morphological and transcriptomic tokens, fuses them via cross-attention, and incorporates a spatially guided token merging mechanism alongside a diffusion-based decoder to enable high-resolution cellular image reconstruction. The framework supports joint representation learning across cellular, microenvironmental, and tissue-level scales. Evaluated on 12 downstream tasks, it consistently outperforms 13 baseline models, achieving significant gains in cell annotation, spatial clustering, gene imputation, and cross-modal prediction. Notably, it is the first method to generate biologically plausible cellular morphologies conditioned on transcriptomic states.
Existing single-cell large language models (LLMs) struggle to effectively integrate spatial coordinates with cell–cell interaction information, resulting in inadequate spatial semantic modeling and insufficient capture of biological relationships. To address this, we propose Spatial2Sentence—a multi-sentence framework tailored for imaging mass cytometry (IMC) data. Our method introduces a novel spatial–expression dual positive/negative sampling paradigm that jointly encodes single-cell expression profiles and spatial proximity into natural language sequences. It is the first to semantically represent spatial coordinates as part of a “cellular language” and explicitly model bidirectional interactions between spatial and functional modalities. The framework integrates multi-task learning, distance-matrix-guided sample construction, and LLM-driven cross-modal textual encoding. Evaluated on a diabetic IMC dataset, Spatial2Sentence achieves absolute improvements of 5.98% in cell-type classification accuracy and 4.18% in clinical state prediction accuracy, while significantly enhancing model interpretability and biological relevance.
This study aims to digitally model tumor microenvironment (TME) heterogeneity by predicting micrometer-scale (55–100 μm) spatial pathway activity directly from routine hematoxylin–eosin (H&E)-stained histopathology images—without requiring sequencing. Method: We introduce the first application of foundational pathology large language models (PLMs) to extract high-fidelity image features, coupled with linear and nonlinear regression for pathway activity prediction. The approach is rigorously validated across three independent spatial transcriptomics (ST) datasets from breast and lung cancers. Contribution/Results: We demonstrate that PLM-derived features accurately recapitulate spatial activity patterns of core signaling pathways—including TGFβ—achieving optimal predictive performance (87–88% reliable samples). Predicted spatial activity maps exhibit clear biological contrast between tumor and non-tumor regions. This establishes a novel computational pathology paradigm for pathway-level, spatially resolved TME characterization directly from standard H&E slides.
Spatial transcriptomics is limited by low experimental throughput and high cost, permitting measurement of only a sparse subset of genes—necessitating reliable imputation of unmeasured gene expression. To address this, we propose a cross-attention–based multimodal deep learning framework that leverages cell-type centroids derived from single-cell RNA-seq to infer spatially resolved gene expression across modalities. Crucially, our method operates without paired samples and explicitly models shared gene co-expression structure between spatial and single-cell modalities to achieve accurate cross-modal alignment and expression recovery. We systematically evaluate the framework on four public spatial transcriptomics datasets. Across 12 quantitative metrics, our approach significantly outperforms state-of-the-art methods in 9, substantially improving prediction accuracy for unmeasured genes. The framework establishes a scalable, interpretable paradigm for multimodal integration in spatial omics, enabling robust inference of spatial gene expression patterns from complementary single-cell data.
Cellular identity and function are linked to both their intrinsic genomic makeup and extrinsic spatial context within the tissue microenvironment. Spatial transcriptomics (ST) offers an unprecedented opportunity to study this, providing in situ gene expression profiles at single-cell resolution and illuminating the spatial and functional organization of cells within tissues. However, a significant hurdle remains: ST data is inherently noisy, large, and structurally complex. This complexity makes it intractable for existing computational methods to effectively capture the interplay between spatial interactions and intrinsic genomic relationships, thus limiting our ability to discern critical biological patterns. Here, we present CellScape, a deep learning framework designed to overcome these limitations for high-performance ST data analysis and pattern discovery. CellScape jointly models cellular interactions in tissue space and genomic relationships among cells, producing comprehensive representations that seamlessly integrate spatial signals with underlying gene regulatory mechanisms. This technique uncovers biologically informative patterns that improve spatial domain segmentation and supports comprehensive spatial cellular analyses across diverse transcriptomics datasets, offering an accurate and versatile framework for deep analysis and interpretation of ST data.w
Existing spatial transcriptomics analysis methods often overlook inter-gene correlations and rely on predefined covariance kernels, which can lead to inflated false positives and false negatives. To address this limitation, this work proposes JASPER—a Bayesian joint modeling framework that, for the first time in spatial transcriptomics, simultaneously captures multi-gene expression patterns through spatial basis function regression without assuming a prespecified covariance structure. By avoiding rigid parametric assumptions about spatial dependence, JASPER substantially enhances statistical robustness and biological interpretability. Extensive evaluation on both real and simulated datasets demonstrates that JASPER consistently identifies gene modules with stronger spatial coherence and clearer functional relevance, as corroborated by pathway and enrichment analyses.
This study addresses the challenge of integrating local microenvironments with global pseudotemporal trajectories in spatial transcriptomics, which involves multimodal registration across samples and regions as well as deciphering complex spatiotemporal expression patterns. To this end, the authors propose a multi-region analytical paradigm that jointly models local neighborhoods and global developmental trajectories. They develop an integrated visual analytics system featuring novel glyphs and a computational framework to enable efficient, interactive exploration of spatial transcriptomic data alongside reference cell atlases and simulated temporal dynamics. In case studies involving pathologists and oncologists, as well as external evaluations, the system effectively facilitated the identification of cellular state transitions and the discovery of spatiotemporal gene expression dynamics.
Spatial transcriptomics remains limited in widespread adoption due to high costs and low throughput, creating an urgent need for methods that can accurately predict gene expression from routine H&E images. To address this challenge, this work proposes COAST, a novel framework that explicitly models relative expression differences between spatial locations. COAST integrates type-specific contextual modulation with a Transformer encoder to jointly capture local fine-grained patterns and whole-slide structural information. A tailored joint loss function is introduced to simultaneously optimize both absolute expression values and signed differential relationships. Evaluated across multiple datasets, COAST significantly improves prediction correlation and distributional consistency, demonstrating the efficacy of context-aware differential learning for spatial gene expression inference.
Spatial transcriptomics (ST) faces challenges in achieving high-throughput, subcellular-scale profiling of thousands of genes while preserving single-cell-resolution spatial context. This work proposes a cross-modal translation approach based on adversarial fine-tuning that leverages unpaired single-cell RNA sequencing (scRNA-seq) and ST data. By fine-tuning a pre-trained single-cell foundation model, the method maps scRNA-seq profiles to spatial coordinates without requiring paired multi-omic samples. It represents the first application of adversarial fine-tuning to single-cell foundation models, thereby overcoming the longstanding dependency on matched datasets for multi-omic integration. The approach significantly outperforms existing methods in reconstructing native cellular neighborhood structures, offering a powerful framework for spatially informed single-cell analysis.
Existing approaches struggle to effectively integrate histological morphology with spatial genomic data and model their spatial context, limiting the accuracy of in silico gene expression prediction from H&E images and its clinical prognostic utility. To address this, this work proposes JASPR—a self-supervised deep learning framework that, for the first time, explicitly aligns the spatial contexts of histology and spatial transcriptomics within a self-supervised paradigm. JASPR employs a shared-expert architecture to jointly learn both cross-modal shared representations and modality-specific features, leveraging a cross-modal reconstruction objective to integrate whole-slide H&E images with spatial transcriptomic data. Evaluated on breast cancer datasets, JASPR significantly improves the accuracy of virtual expression prediction for 9,248 genes and demonstrates independent clinical prognostic value.