gene-level perturbation simulation

Designs and implements predictive and generative models that simulate single-cell transcriptomic responses to gene- or drug-level perturbations by conditioning on cell state and treatment, producing post‑treatment expression profiles and transferable perturbation representations. These systems include attention-based gene–gene or multimodal architectures that generalize to unseen genes or treatments and incorporate cell‑cycle–aware mechanisms (circular‑phase heads, closed‑loop supervision) that predict or enforce treatment‑induced cell‑cycle changes and propagate supervision gradients back to expression.

gene-levelperturbationsimulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing single-cell drug perturbation models struggle to accurately predict the effects of drugs on cell cycle phase transitions, as they typically treat the cell cycle as a nuisance factor rather than explicitly modeling it. This work proposes scCycleMol, a novel framework that infers cell cycle supervision signals from post-treatment gene expression profiles in the SciPlex3 dataset. By reframing cell cycle state as a learnable output target instead of an input covariate, scCycleMol establishes a closed-loop supervision mechanism. Integrating molecular representations, pretraining strategies, circular modeling of G1/S/G2M phases, and a conditional perturbation architecture, the method achieves state-of-the-art performance across over 600,000 cells, yielding an R² of 0.9093 for all genes, 0.6843 for differentially expressed genes, and a phase classification accuracy of 0.9609—significantly outperforming baselines such as ChemCPA.

cell-cycle statephase predictionproliferative state

Modeling Gene Expression Distributional Shifts for Unseen Genetic Perturbations

Jul 01, 2025
KR
Kalyan Ramakrishnan
🏛️ University of Oxford | Novo Nordisk

In early drug discovery, existing gene perturbation prediction methods model only mean expression levels, failing to capture cellular heterogeneity. This work introduces the first deep learning framework capable of predicting the full single-cell gene expression distribution—including variance, skewness, and kurtosis. Methodologically, it innovatively adopts gene-level histograms as output targets and integrates large language model–derived gene embeddings as biologically informed priors to enable generalization to unseen perturbations. Experiments demonstrate that our model significantly outperforms baselines in distributional modeling (−12.7% KL divergence), reduces training cost by 35%, and maintains state-of-the-art accuracy in mean expression prediction. By enabling high-fidelity, distribution-aware perturbation response modeling, this work establishes a more realistic and robust paradigm for target identification and functional interpretation in perturbation biology.

Generalize to unseen perturbations using gene embeddings from LLMsOvercome limitations of mean-only prediction in single-cell dataPredict gene expression distribution shifts post genetic perturbations

This study addresses the challenging problem of predicting transcriptional responses to unseen gene perturbations in single cells under a zero-shot setting, where perturbation effects arise from complex interactions among cellular states, gene functions, and regulatory mechanisms. To tackle this, the authors propose CisTransCell, a novel framework that, for the first time, incorporates both cis-regulatory sequences and trans-encoded protein sequences as multimodal priors alongside cellular expression states. CisTransCell models the cascade from gene function through regulatory logic to downstream transcriptional changes. By integrating genomic, proteomic coding, and single-cell transcriptomic data, the method substantially outperforms existing approaches across multiple benchmark datasets, demonstrating exceptional performance in zero-shot gene perturbation prediction.

cellular contextgene regulationsingle-cell perturbation prediction

This work addresses the challenge in single-cell perturbation prediction that perturbation-specific signals are sparse and often obscured by dominant invariant expression structures, hindering existing methods from learning generalizable causal representations. To overcome this, the authors propose PerturbedVAE, a novel framework that explicitly disentangles perturbation-specific and invariant features for the first time, guided by identifiability theory to recover sparse perturbation effects. Built upon a variational autoencoder architecture, PerturbedVAE supports modeling of combinatorial perturbations and achieves substantially improved out-of-distribution prediction performance on standard benchmarks. Furthermore, the model reveals interpretable perturbation–response mechanisms, offering new insights into cellular response dynamics.

invariant informationperturbation-specific signalsrepresentation learning

This study addresses the neglect of temporal gene response dynamics in existing single-cell perturbation prediction methods by proposing the D²R² framework. This approach reformulates prediction as a regulation-guided, progressive gene-by-generation process, integrating masked discrete diffusion models with reinforcement learning (GRPO). Notably, it is the first to treat generation order as an interpretable dimension subject to joint optimization. Experiments on the Norman19 dataset demonstrate state-of-the-art performance across five metrics with strong competitiveness. Furthermore, ablation studies confirm that biologically informed ordering significantly outperforms baseline strategies, and the optimized generation paths accurately capture key regulatory factors, thereby achieving a unification of high predictive accuracy and model interpretability.

Functional GenomicsGene Expression ProfileGeneration Order

Latest Papers

What's happening recently
View more

This study addresses the limited predictive and interpretive capacity of existing models, which lack explicit representations of cellular state transition mechanisms under genetic perturbations. We propose a mechanism-centric virtual cell world model that represents cellular states as Latent Mechanistic Units (LMUs) and treats genetic perturbations as actions operating on these units, thereby simulating mechanism-level stochastic state transitions. By integrating a reusable identity module with observation-specific state architectures grounded in multimodal evidence, our approach transcends conventional direct-mapping paradigms. Training follows a two-stage strategy comprising large-scale pseudo-bulk perturbation profile pretraining followed by single-cell data fine-tuning. Evaluated across six disjoint benchmarks, the proposed model significantly improves the accuracy of perturbation-specific response recovery, demonstrating the capacity of LMUs to capture structured biological programs.

cellular state transitiongenetic perturbationperturbation response

This study addresses the limitations of existing cellular perturbation prediction methods in achieving continuous dose control and morphological transformation. We propose a dual-time-step joint flow matching framework that simultaneously models cellular latent variables and drug concentrations. Leveraging the invertibility of flow matching, this approach enables single-cell morphology generation under continuous dosing and inverse estimation of concentrations from observed morphologies. Experimental results demonstrate that the proposed model matches or exceeds baseline performance on two compounds and successfully generalizes to unseen doses. Consequently, this work provides an effective generative solution for modeling continuous drug responses, bridging the gap between discrete experimental observations and continuous pharmacological dynamics in single-cell perturbation studies.

Cellular perturbation predictionContinuous dose controlDose-conditioned cell morphing

This study addresses the challenge of reverse-identifying molecular targets and compounds from single-cell transcriptional perturbation responses. The authors propose the first multi-task Transformer-based retrieval model tailored for single-cell perturbation data, which jointly learns target prediction and molecular embedding within a fixed compound library. The model performs end-to-end inverse inference using differential expression profiles relative to cell-type-specific DMSO controls and incorporates a structure–transcriptome alignment constraint to enhance representational consistency. Evaluated on the Tahoe-100M dataset, the model achieves a target Recall@10 of 0.408 and a compound Hit@1 of 0.129, significantly outperforming baseline methods and demonstrating its effectiveness in retrieving known perturbation pairs.

compound retrievalinverse problemperturbation signature

Hot Scholars

DZ

Ding Zou

Huazhong University of Science and Technology
RecommendationKnowledge graphContrastive Learning
SZ

Sen Zhao

Chongqing University of Posts and Telecommunications
Recommendation systemsInformation RetrievalNature language processing
AM

Anil Madhavapeddy

Professor of Planetary Computing, University of Cambridge
Computer Science
JC

Jon Crowcroft

University of Cambridge
CommunicationsSystems
AS

Abhronil Sengupta

Monkowski Career Development Associate Professor of EECS, Penn State University
Neuromorphic Computing