🤖 AI Summary
This study addresses the challenges of scale mismatch, context loss, and computational inefficiency in predicting single-cell gene expression from H&E images by proposing CELLO. This framework introduces the first end-to-end architecture for single-cell prediction, leveraging a single forward pass through a pathology foundation model combined with grid sampling. By incorporating distance-decayed cross-attention, CELLO efficiently extracts context-aware local morphological features without requiring per-cell cropping, thereby effectively avoiding segmentation error accumulation. Experimental evaluations across 52 datasets demonstrate that CELLO significantly improves prediction accuracy while achieving an average 14-fold speedup over baseline methods. Ultimately, this work establishes a new paradigm for scalable single-cell spatial transcriptomics prediction.
📝 Abstract
Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-H&E pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from H&E images.