CellMSA: Context Modeling for Single-Cell Representation Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that high dimensionality, sparsity, and batch effects in single-cell transcriptomic data cause models to overlook cross-batch associations. Inspired by protein multiple sequence alignment (MSA), we propose a single-cell representation learning framework incorporating an MSA inductive bias. The method employs a context retrieval mechanism to extract cross-batch correlated cells for context construction, and designs a pair-aware encoder to generate gene-pair representations, capturing fine-grained gene dependencies via cross-cell consistency and overcoming the limitations of traditional independent encoding. The model is pretrained on a large-scale corpus of 109 million cells. Extensive benchmark evaluations demonstrate that it significantly outperforms existing methods, achieving efficient cross-batch data integration.
📝 Abstract
Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells from the same batch for denoising, thereby underutilizing the rich relational information across batches and cell types to model gene expression patterns. We argue that single-cell models can benefit from more informative cell-context modeling. By comparing consistency and variation across cells, models can capture fine-grained gene-gene dependencies associated with cell states, which are essential for learning high-quality representations. Inspired by the use of multiple sequence alignment (MSA) context in protein modeling, we propose CellMSA, a single-cell representation learning framework that introduces an MSA-inspired inductive bias into transcriptomic modeling. For each target cell, CellMSA retrieves relevant cells from different batches and biologically related cell types as context, and summarizes cross-cell patterns into a context-dependent gene-pair representation. This representation is then injected into a pair-aware target-cell encoder for fine-grained representation learning. We pretrain CellMSA on a large-scale human single-cell corpus of approximately 109 million cell observations, including 65.6 million primary observations. Experiments show that our framework consistently outperforms existing methods across multiple benchmarks. Code is available at the following repository: https://github.com/PharMolix/CellMSA.
Problem

Research questions and friction points this paper is trying to address.

single-cell transcriptomics
representation learning
context modeling
batch effects
gene-gene dependencies
Innovation

Methods, ideas, or system contributions that make the work stand out.

CellMSA
Multiple Sequence Alignment
Context Modeling
Single-Cell Representation Learning
Pair-aware Encoder
🔎 Similar Papers
2024-08-22Neural Information Processing SystemsCitations: 0
💼 Related Jobs
No related jobs found.
S
Suyuan Zhao
Institute for AI Industry Research (AIR), Tsinghua University; Department of Computer Science and Technology, Tsinghua University
M
Minghao Liu
Institute for AI Industry Research (AIR), Tsinghua University; Tsinghua Institute of Multidisciplinary Biomedical Research (TIMBR), Tsinghua University; National Institute of Biological Sciences (NIBS)
Y
Yizhen Luo
Institute for AI Industry Research (AIR), Tsinghua University; Department of Computer Science and Technology, Tsinghua University
Zaiqing Nie
Zaiqing Nie
Tsinghua University
NLPData MiningMachine Learning