Prototype Guided Post-pretraining for Single-Cell Representation Learning

📅 2026-05-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Single-cell pre-trained models often suffer from limited generalization due to the long-tailed distribution of cell types and covariate shifts in gene expression data. To address this, this work proposes CellRefine, a novel framework that introduces a prototype-guided post-pretraining phase between pretraining and fine-tuning. For the first time, CellRefine incorporates curated marker gene sets as biologically informed structural priors and refines the latent cell embedding manifold through multi-objective optimization. This approach substantially enhances the generalization performance of single-cell foundation models across diverse downstream tasks, achieving performance gains of up to 15%.
📝 Abstract
Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular function. Inspired by large language models in natural language modeling, several single-cell pretrained models have recently been proposed that treat genes as tokens and cells as sentences. However, these models are fundamentally limited by the long-tailed nature of cell-type distributions and struggle to generalize under covariate shifts in gene expression data. While fine-tuning is often used to mitigate these issues, we observe that performance remains bounded. To address this challenge, we introduce CellRefine, a post-pretraining method that operates between the pretraining and fine-tuning stages of a single-cell foundation model. CellRefine uses a multi-faceted objective that incorporates marker-gene sets as structural priors to guide post-pretraining and refine the latent embedding manifold of cells. Across multiple computational biology tasks, empirical results show that CellRefine consistently improves downstream performance, yielding gains up to 15%.
Problem

Research questions and friction points this paper is trying to address.

single-cell representation learning
long-tailed distribution
covariate shift
pretrained models
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-pretraining
single-cell representation learning
marker-gene priors
embedding manifold refinement
foundation model
🔎 Similar Papers
2024-08-22Neural Information Processing SystemsCitations: 0
S
Sachini Weerasekara
N
Natasha Darras
S
Sagar Kamarthi
C
Colles Price
J
Jacqueline Isaacs