single-cell data analysis

Computational methods for processing and interpreting single-cell and single-cell perturbation data, including embedding construction, latent recovery, network reconstruction, and imputation, and empirical benchmarking against manifold-preserving methods under realistic sampling conditions.

single-celldataanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge in single-cell perturbation prediction that perturbation-specific signals are sparse and often obscured by dominant invariant expression structures, hindering existing methods from learning generalizable causal representations. To overcome this, the authors propose PerturbedVAE, a novel framework that explicitly disentangles perturbation-specific and invariant features for the first time, guided by identifiability theory to recover sparse perturbation effects. Built upon a variational autoencoder architecture, PerturbedVAE supports modeling of combinatorial perturbations and achieves substantially improved out-of-distribution prediction performance on standard benchmarks. Furthermore, the model reveals interpretable perturbation–response mechanisms, offering new insights into cellular response dynamics.

invariant informationperturbation-specific signalsrepresentation learning

Lack of standardized evaluation protocols for single-cell perturbation effect prediction hinders model comparability and biological interpretability. To address this, we introduce ScPEBench—the first standardized benchmark framework—integrating diverse CRISPR-based and drug perturbation datasets, an accessible benchmarking platform, and a multi-dimensional evaluation suite balancing reconstruction accuracy (e.g., RMSE) and ranking fidelity (e.g., rank-based consistency). Through systematic, reproducible integration testing, we uncover pervasive issues—including mode collapse—and demonstrate that simple models often outperform complex ones, underscoring the critical importance of ranking metrics alongside conventional error measures. ScPEBench establishes a new evaluation paradigm, significantly enhancing model robustness and cross-study comparability. It provides a rigorous, reproducible foundation for genetic and chemical screening–driven target discovery.

Advancing disease target discovery via high-throughput perturbation screensComparing model performance fairly using diverse datasets and metricsStandardizing benchmarking for cellular perturbation response modeling

This work addresses the challenge that in single-cell perturbation data, cell populations under different perturbations exhibit substantial overlap, rendering conventional single-cell classification accuracy an unreliable metric of model performance. To overcome this limitation, the authors propose the Classifier Discrimination Score (CDS), which constructs a perturbation-level profile by aggregating classifier output probability distributions across entire cell populations and replaces single-cell predictions with population-level ranking. Remarkably, CDS recovers near-perfect perturbation identification from weak classifiers without requiring retraining. The method is compatible with diverse architectures—including linear models, MLPs, and Transformers—and demonstrates significant gains in identification accuracy on the Tahoe-100M and Virtual Cell Challenge datasets, with particularly pronounced advantages in low-cell-count regimes.

class overlapclassification accuracymodel evaluation

This study addresses the challenge of limited observational data for specific perturbations in target cellular contexts within single-cell perturbation atlases, where direct transfer of responses from source environments often introduces bias. To mitigate this, the authors propose a path reliability–weighted transfer mechanism that integrates a local low-rank basis from the target environment with a source-to-target ridge regression expert network. The reliability of transfer paths is evaluated using validation anchors, and existing perturbation responses are retrieved and aggregated via cosine similarity–based weighting. This approach effectively incorporates cross-environment information while preserving perturbation specificity. Evaluated on the Perturb-CITE-seq melanoma dataset, the method reduces mean squared error by 4.1% and improves top-10 in-context perturbation retrieval accuracy from 74.5% to 80.5% compared to the local low-rank baseline.

cellular contextcross-context transfermissing perturbation response

This work addresses the fundamental challenge in single-cell perturbation experiments where the destructive nature of sequencing precludes observing pre- and post-perturbation states in the same cell, and existing methods struggle to model multimodal response distributions arising from latent variables such as microenvironmental context and batch effects. To overcome these limitations, the authors introduce— for the first time—a diffusion generative model operating directly in distribution space. By embedding cellular population distributions into a reproducing kernel Hilbert space (RKHS), they construct a diffusion process that acts on probability measures, explicitly modeling the distributional evolution induced by perturbations. This approach moves beyond the conventional single-response assumption and achieves state-of-the-art performance across multiple single-cell transcriptomic benchmark datasets, substantially improving generalization to unseen perturbations.

distribution shiftlatent factorsresponse variability

Latest Papers

What's happening recently
View more

Predicting transcriptional responses of cells to perturbations is challenging due to the high noise and sparsity of single-cell data, as well as the fact that perturbations typically induce shifts in population-level distributions rather than deterministic changes at the single-cell level. This work proposes the first distribution-level generative modeling framework for single-cell perturbation prediction, leveraging conditional flow matching to learn the full post-perturbation cell population distribution. By incorporating Maximum Mean Discrepancy (MMD) to align control and perturbed populations, the method eliminates reliance on one-to-one single-cell correspondences. Central to this approach is the Perturbation-Aware Differential Transformer (PAD-Transformer), which integrates gene interaction graph attention to effectively capture context-specific expression changes. The model significantly outperforms existing methods across diverse genetic and drug perturbation benchmarks, achieving a 19.6% reduction in mean squared error over the strongest baseline in combinatorial perturbation settings, demonstrating exceptional generalization capability.

distributional shiftnoisy and sparse datapopulation-level effects

This work addresses the challenge in single-cell perturbation prediction of simultaneously modeling the implicit mechanisms of perturbations and their temporal dynamics, which limits generalization to unseen interventions. The authors propose CITE-VAE, an implicit dynamical causal generative model that formalizes perturbation effects as causal mechanisms exhibiting both latent structure and temporal evolution. CITE-VAE jointly models latent cellular programs, perturbation conditions, and time-dependent dynamics, and provides theoretical identifiability guarantees for latent variables within standard equivalence classes. Efficient inference is achieved through a variational autoencoder framework. Experiments demonstrate that the model validates its identifiability on the Causal-3DIdent synthetic benchmark and significantly outperforms existing methods on real CRISPR-based single-cell perturbation data, achieving superior out-of-distribution generalization performance.

dynamical systemslatent causal processesout-of-distribution generalization

This work addresses the challenge posed by pervasive dropout noise—exceeding 90% in single-cell transcriptomic data—which causes existing models to learn technical artifacts rather than stable biological programs under reconstruction objectives. To overcome this, the authors propose Cell-JEPA, the first method to adapt the Joint Embedding Predictive Architecture (JEPA) to single-cell modeling. By predicting complete cell embeddings from partially observed inputs in a latent space, Cell-JEPA avoids direct reconstruction of sparse, noisy expression counts and instead leverages gene redundancy to learn representations robust to dropout. The model achieves an AvgBIO score of 0.72 on zero-shot cell-type clustering, a 36% improvement over scGPT, and enhances the accuracy of cellular state reconstruction in perturbation response prediction tasks.

dropoutfoundation modelsrepresentation learning

This study addresses the challenge of reverse-identifying molecular targets and compounds from single-cell transcriptional perturbation responses. The authors propose the first multi-task Transformer-based retrieval model tailored for single-cell perturbation data, which jointly learns target prediction and molecular embedding within a fixed compound library. The model performs end-to-end inverse inference using differential expression profiles relative to cell-type-specific DMSO controls and incorporates a structure–transcriptome alignment constraint to enhance representational consistency. Evaluated on the Tahoe-100M dataset, the model achieves a target Recall@10 of 0.408 and a compound Hit@1 of 0.129, significantly outperforming baseline methods and demonstrating its effectiveness in retrieving known perturbation pairs.

compound retrievalinverse problemperturbation signature

This work addresses the high storage and reuse costs of single-cell data and the lack of auditability in existing distillation methods that produce non-traceable synthetic data. The authors propose Minmax-CF, a method that, under fixed budgets of cells and genes, constructs a traceable core set by selecting only real observed cells through discrete minimax optimization and matching with static feature functions. This approach fully preserves original cell identifiers, gene symbols, and associated counts, labels, and metadata. Evaluated across multiple datasets, Minmax-CF achieves up to 96.52% balanced accuracy, substantially reduces pathway errors, and yields up to 2.55× GPU acceleration. The selected cells further enable efficient anomaly analysis and model validation.

auditabledata provenancereal-cell coresets

Hot Scholars

YH

Yuankai Huo

Computer Science, Vanderbilt University
Medical Image AnalysisDeep LearningData Mining
YZ

Yuanchun Zhou

Computer Network Information Center,CAS
Data MiningBig Data Analysis
SK

Smita Krishnaswamy

Yale University
Machine LearningData MiningManifold LearningDeep Learning
RD

Ruining Deng

Weill Cornell Medicine
Medical Image AnalysisDeep LearningDigital Pathology