differential expression analysis

Statistical and machine‑learning methods to detect and predict changes in gene or cell‑level expression across conditions, design biologically meaningful verifiers/reward functions for single‑cell perturbations, and evaluate generalization across datasets and cell lines.

differentialexpressionanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Modeling Gene Expression Distributional Shifts for Unseen Genetic Perturbations

Jul 01, 2025
KR
Kalyan Ramakrishnan
🏛️ University of Oxford | Novo Nordisk

In early drug discovery, existing gene perturbation prediction methods model only mean expression levels, failing to capture cellular heterogeneity. This work introduces the first deep learning framework capable of predicting the full single-cell gene expression distribution—including variance, skewness, and kurtosis. Methodologically, it innovatively adopts gene-level histograms as output targets and integrates large language model–derived gene embeddings as biologically informed priors to enable generalization to unseen perturbations. Experiments demonstrate that our model significantly outperforms baselines in distributional modeling (−12.7% KL divergence), reduces training cost by 35%, and maintains state-of-the-art accuracy in mean expression prediction. By enabling high-fidelity, distribution-aware perturbation response modeling, this work establishes a more realistic and robust paradigm for target identification and functional interpretation in perturbation biology.

Generalize to unseen perturbations using gene embeddings from LLMsOvercome limitations of mean-only prediction in single-cell dataPredict gene expression distribution shifts post genetic perturbations

Automated Statistical and Machine Learning Platform for Biological Research

Nov 25, 2025
LR
Luke Rimmo Lego
🏛️ Stevens Institute of Technology

In biological research, the fragmentation between statistical analysis and machine learning tools, coupled with high usability barriers for non-programming users, impedes efficient and rigorous data-driven discovery. Method: We propose BioAutoML, a modular, biology-oriented automated analysis platform integrating classical statistical methods (e.g., t-tests, ANOVA, Pearson correlation) with interpretable machine learning (e.g., Random Forest classification). It supports automated data preprocessing, categorical encoding, feature importance assessment, and data-aware dynamic model configuration. Crucially, it introduces the first unified statistical–machine learning workflow, bridging methodological gaps via automated hyperparameter optimization. Contribution/Results: Evaluated on multiple chemomics datasets, BioAutoML achieves significantly higher classification accuracy than baseline approaches while preserving statistical validity. It enables domain scientists without programming expertise to perform end-to-end, interpretable, and statistically sound modeling—substantially accelerating biological insight generation.

Automates hyperparameter optimization and feature importance for non-programmersIntegrates statistical and machine learning methods for biological data analysisUnifies diverse tools to streamline workflows and enhance interpretability in bioinformatics

This study addresses the challenge of effectively integrating heterogeneous proteomic data—specifically, whole-sample mass spectrometry (MS) and multiplexed protein array profiles—in pancreatic cancer research. To overcome the limitations of conventional approaches that naively concatenate multi-source features, the authors propose a novel model fusion framework that explicitly models and leverages the heterogeneity between data sources through a tailored integration strategy. This approach synergistically exploits the complementary strengths of each modality rather than treating them as homogeneous inputs. Experimental results demonstrate that the proposed method significantly outperforms both single-modality models and standard fusion baselines in pancreatic cancer classification, yielding substantial improvements in diagnostic accuracy. The work thus offers a principled and effective paradigm for integrating heterogeneous multi-omics data in biomedical applications.

classification modelsdata integrationheterogeneous proteomic data

Single-cell RNA-seq data are inherently noisy and sparse, and existing visualization methods often lack a solid statistical foundation, making it difficult to distinguish technical artifacts from genuine biological signals. This work proposes ZINBGT—a zero-inflated negative binomial mixture model with a geometric tail—that uniquely integrates statistical rigor with biological interpretability. The model provides interpretable visualizations of gene expression at the individual gene level and employs Wasserstein distance to diagnose aberrant genes. Applying ZINBGT, the authors successfully identify outlier-expressing genes in *T. brucei*, uncover intrinsic relationships among sparsity, mean expression, and dispersion in human immune cells, and highlight limitations of current simulated datasets in accurately recapitulating key characteristics of real single-cell data.

data noisesingle-cell transcriptomicsstatistical inference

Data-Driven Logistic Regression Ensembles With Applications in Genomics

Feb 17, 2021
AC
A. Christidis
🏛️ University of British Columbia | KU Leuven

Addressing the challenge of simultaneously achieving high prediction accuracy and biologically interpretable biomarker identification in high-dimensional genomic binary classification, this paper proposes the first data-driven logistic regression (LR) ensemble framework that jointly ensures statistical interpretability and strong predictive performance. The method integrates L₁/L₂ regularization with ensemble learning to directly learn a small set of highly accurate and interpretable LR models via global optimization. It provides the first rigorous derivation of asymptotic statistical properties for such regularized ensemble LR estimators. Additionally, we develop a gene importance ranking tool based on resampling stability to enhance biological interpretability. Evaluated on real-world datasets—including cancer, multiple sclerosis, and psoriasis—the framework achieves significant improvements in classification accuracy and successfully identifies several critical disease-associated genes—previously missed by competing methods—that are independently validated in the biomedical literature.

Develops data-driven logistic regression ensembles for genomicsImproves prediction accuracy and biomarker identification in diseasesProvides variable importance ranking for prioritizing critical genes

Latest Papers

What's happening recently
View more

The molecular mechanisms underlying multiple sclerosis (MS) remain incompletely understood. This study develops an end-to-end machine learning pipeline that integrates bulk microarray and single-cell RNA-seq data from peripheral blood and cerebrospinal fluid (CSF) to distinguish MS patients from healthy controls using an XGBoost classifier. By combining SHAP-based interpretability with differential expression analysis, the work uncovers novel mechanisms at the multi-tissue transcriptomic level, including non-canonical immune checkpoints and virus-related pathways. The model achieves high performance in CSF B cells (AUC = 0.94) and microarray data (AUC = 0.86), identifying several candidate biomarkers linked to immune activation, the ubiquitin–proteasome system, and Epstein–Barr virus infection.

biomarker discoverycross-tissue analysisimmune mechanisms

Current single-cell perturbation prediction models lack explicit constraints to ensure biological consistency when generating individual cells, leading to unreliable outputs. This work proposes a reinforcement learning–based post-training framework that, for the first time, incorporates multidimensional cell-level validators—including Pearson top-k similarity, RMSE top-k proximity, Spearman correlation of differentially expressed genes, and pathway activity—as reward signals to guide a pretrained generator toward biologically plausible perturbation responses. The method significantly improves alignment with these validators and enhances held-out evaluation metrics across multiple genetic and chemical perturbation benchmarks, while maintaining population-level performance comparable to state-of-the-art approaches, thereby advancing a more trustworthy paradigm for single-cell perturbation prediction.

biological consistencygenerative modelssingle-cell perturbation

Current large language models struggle to accurately predict gene expression changes under specific cellular perturbations, often conflating intrinsic gene responses with genuine perturbation effects. To address this, this work proposes the Contrastive Relational Evidence Organization (CORE) framework, which reframes perturbation prediction as a contrastive task for the first time. CORE leverages biomedical knowledge graphs to retrieve positive and negative regulatory effects of the same gene across different perturbations, thereby enhancing the model’s ability to reason about perturbation-specific responses. The framework supports two paradigms—CORE-Reasoning and CORE-Voting—and demonstrates substantial improvements: on drug perturbation data, it boosts the aggregate metric of Qwen3.5-9B by up to 28.6%; on general perturbation benchmarks, CORE-Voting elevates the average macro-gene AUROC from random chance to 0.703, underscoring the critical role of relational evidence in improving both prediction accuracy and calibration.

cellular perturbationcontrastive evidencedifferential expression

This study addresses the challenge of subtype-specific diagnosis in non-small cell lung cancer (NSCLC) by proposing a two-stage approach: first, integrating differential expression and methylation analyses to identify adenocarcinoma (LUAD) and squamous cell carcinoma (LUSC)-specific biomarkers, and then constructing a quantum machine learning classifier to accurately distinguish among LUAD, LUSC, and normal samples. This work represents the first application of quantum computing to multi-omics analysis in lung cancer, uncovering key genes—including NGFR, NTRK2, and NTF3—and implicating neurotrophin, MAPK, Ras, and PI3K-Akt signaling pathways in tumorigenesis, thereby highlighting the role of neural signaling in cancer development. The quantum classifier demonstrates superior performance and scalability on biomedical big data, with the Sample3 gene set achieving optimal results across all evaluation metrics.

biomarker discoverycancer diagnosticslung cancer subtypes

This study addresses the limited interpretability of existing rare cell identification methods in single-cell transcriptomics, which typically rely on opaque dimensionality reduction techniques (e.g., PCA) and black-box anomaly detectors, offering little insight at the gene level. To overcome this, we propose an end-to-end, interpretable anomaly detection framework that operates directly in the original high-dimensional gene expression space without requiring dimensionality reduction. Our method jointly optimizes anomaly detection and feature attribution, enabling precise identification of rare cells while simultaneously highlighting their key discriminative genes. By integrating state-of-the-art interpretable anomaly detection into single-cell analysis for the first time, our approach not only locates rare cells but also links them to their nearest normal neighbors, providing intuitive, biologically meaningful gene-level explanations that substantially enhance both interpretability and practical utility.

anomaly detectionexplainable AIinterpretability

Hot Scholars

YZ

Yuanchun Zhou

Computer Network Information Center,CAS
Data MiningBig Data Analysis
HW

Haohan Wang

School of Information Sciences, University of Illinois Urbana-Champaign
Computational BiologyAgentic AIAI4ScienceAI security
JH

Junzhou Huang

Jenkins Garrett Professor, Computer Science and Engineering, the University of Texas at Arlington
Machine LearningMedical Image AnalysisGraph Neural NetworksComputational Toxicology
SL

Sidong Liu

Australian Institute of Health Innovation, Macquarie University
Medical Image ComputingComputational NeurosciencePersonalized Oncology
YM

Yuwei Miao

PhD student, University of Texas at Arlington