gene regulatory network inference

Designs and implements methods to reconstruct directed networks of regulatory interactions among genes, inferring which genes regulate which others and estimating the strengths and signs of those influences from observational molecular data. This includes approaches that infer networks from snapshot or time‑series measurements, align gene identities across time or cells, and denoise or integrate single‑cell and bulk data to improve network estimates.

generegulatorynetworkinference

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of spurious edge generation when reconstructing causal regulatory networks of dynamical systems from discrete-state trajectories. The authors propose a modeling framework based on integral-form additive nonparametric ordinary differential equations, which accommodates both dense and sparse irregular sampling. By incorporating a data-driven edge selection mechanism, the method effectively suppresses false connections while inferring time-varying, weighted, bidirectional causal relationships between nodes, including activation or inhibition effects. This enables the construction of interpretable, symbolic causal networks. In five simulation experiments, the approach substantially outperforms GRADE, reducing spurious edges from 239 to zero in the most challenging scenario while nearly preserving all true regulatory links, thereby achieving high-precision dynamic network reconstruction.

causal network reconstructiondynamical systemsregulatory network

Machine Learning Methods for Gene Regulatory Network Inference

Apr 17, 2025
AH
Akshata Hegde
🏛️ University of Missouri

This study systematically reviews machine learning—particularly deep learning—for gene regulatory network (GRN) inference, addressing the challenge of modeling nonlinear and dynamic regulatory interactions from high-throughput transcriptomic data (including bulk and single-cell RNA-seq). We propose, for the first time, a unified methodological taxonomy encompassing supervised, unsupervised, semi-supervised, and contrastive learning paradigms. We standardize benchmark datasets (e.g., DREAM, SINCERITIES) and evaluation metrics (AUPR, AUROC, F-score), and empirically delineate current model performance limits. Key contributions include: (1) synthesizing advances in deep neural architectures to elucidate their superior capacity for capturing complex, context-dependent regulatory logic; (2) establishing a reproducible, standardized benchmarking framework with practical implementation guidelines; and (3) laying a methodological foundation for both next-generation GRN algorithm development and mechanism-driven biological discovery.

Analyzing large-scale omics data for gene interactionsEnhancing GRN inference with deep learning techniquesInferring gene regulatory networks using machine learning

Despite theoretical advantages, causal methods for Gene Regulatory Network (GRN) inference from single-cell RNA-seq data consistently fail to match or outperform correlation-based baselines in many realistic benchmarks, a persistent puzzle which casts doubt on the value of causality for this task. We argue that existing benchmarks are insufficiently controlled to answer this question because they evaluate on real or semi-real data where multiple pathologies co-occur, confounding failure modes, and obscuring the specific conditions under which different inference methods excel or fail. To address this gap, we introduce a controlled diagnostic framework that isolates seven biologically motivated pathologies (dropout, latent confounders, cell-type mixing, feedback loops, network density, sample size, and pseudotime drift) and measure how six representative methods spanning three inference paradigms degrade as each pathology intensifies. Across 6,120 controlled experiments, we find that causal methods genuinely dominate in clean and structurally favorable regimes, but specific pathologies (notably dropout and latent confounders) selectively neutralize their advantages. We further introduce an error-type decomposition that reveals methods with similar aggregate accuracy commit qualitatively different errors. To probe whether single-pathology effects persist when multiple stressors co-occur, we perform an interaction sweep over the three most impactful pathologies and find that their joint effects are sub-additive, while also exposing density-conditional cross-overs invisible to single-dial analysis. Our findings offer a nuanced understanding of when and why different methods succeed or fail for GRN inference, providing actionable insights for method development and practical guidance for practitioners.

benchmarkingcausal inferenceGene Regulatory Network

Identifying perturbation targets through causal differential networks

Oct 04, 2024
MW
Menghua Wu
🏛️ Massachusetts Institute of Technology

Identifying intervention targets in single-cell biology—i.e., inferring the set of perturbed variables from combined observational and interventional data—remains challenging due to small sample sizes, high dimensionality, and violations of ideal causal assumptions. Method: We propose Causal Differential Networks (CDN), a novel end-to-end joint training framework that unifies noisy causal graph inference, graph-structural difference modeling, and multi-source feature supervision to jointly optimize for causal interpretability and prediction robustness. Contribution/Results: Evaluated on seven real single-cell transcriptomic datasets and diverse synthetic intervention scenarios, CDN consistently outperforms state-of-the-art baselines. It achieves substantial improvements in both soft and hard target prediction accuracy, offering an interpretable, high-precision computational paradigm for drug target discovery and cellular engineering.

Enhance prediction of soft and hard intervention effects.Identify intervention targets in biological systems.Improve causal discovery in high-dimensional biological data.

This work addresses the challenge of causal discovery in the presence of sample dependencies and mixed continuous-discrete variables by proposing a decorrelation framework based on a latent-variable structural equation model. The method employs an EM algorithm to impute latent variables corresponding to discrete observations and integrates pairwise maximum likelihood covariance estimation with a decorrelation transformation to map the original data into independent and identically distributed latent representations compatible with standard DAG learning algorithms. To the best of our knowledge, this is the first approach that jointly handles mixed variable types and inter-sample dependencies, enabling unified modeling of correlated Gaussian errors and discrete measurements. Experiments demonstrate that the proposed method significantly outperforms existing baselines in simulations and yields gene regulatory networks from single-cell RNA-seq data with higher predictive likelihood, where high-confidence edges show strong concordance with known biological pathways.

causal discoverydependent datagene regulatory network

Latest Papers

What's happening recently
View more

This study investigates the feasibility of recovering causal relationships among genes from aggregated bulk gene expression data. By formalizing the notion of causal recoverability, the authors introduce two key criteria—functional form consistency and conditional independence consistency—and rigorously establish their necessary and sufficient conditions: accurate causal recovery is possible only when gene regulation follows linear aggregation and affine structural equation models. The theoretical analysis integrates causal inference with a functional consistency framework and is empirically validated on multiple bulk and single-cell datasets. Results demonstrate that real gene regulatory mechanisms are predominantly nonlinear, thereby revealing a fundamental limitation of current causal inference methods that rely solely on bulk expression data.

aggregationbulk gene expressioncausal relations

This study addresses key limitations in gene regulatory network (GRN) reconstruction—namely, the tight coupling between modeling assumptions and inference methods, the lack of uncertainty quantification, and the difficulty of transferring prior knowledge across species. To overcome these challenges, the authors propose a two-stage framework: first, a transferable prior model, GLM-Prior, is built by fine-tuning the Nucleotide Transformer on DNA sequences to predict transcription factor–target interactions; second, a probabilistic matrix factorization model, PMF-GRN, incorporates variational inference to enable uncertainty-aware network refinement. This approach uniquely integrates transferable deep sequence modeling with probabilistic inference that quantifies uncertainty, significantly improving the accuracy, generalizability, and reliability of GRN reconstruction across yeast, mouse, and human datasets, while enabling robust evaluation even when reference networks are incomplete.

cross-species generalizationgene regulatory networknetwork inference

This study addresses the challenge of accurately predicting single-cell transcriptional responses to unseen genetic perturbations and drug combinations while avoiding the misattribution of stable gene co-expression patterns as perturbation-induced effects. To this end, the authors propose GeneGeoFlow, a novel method that incorporates perturbation-conditioned, multi-scale gene geometric structures as priors. It employs a perturbation-gated module to dynamically select relevant structural information and constructs a residual flow model anchored on control samples to disentangle stable gene associations from intervention-specific responses. Integrating multi-scale spectral coordinates, anchor-based residual flows, unpaired optimal transport training, and a Delta-correlation objective function, GeneGeoFlow achieves a Pearson Delta score of 0.8979 on the Norman additive benchmark and 0.9088 across five held-out drug combinations in ComboSciPlex, substantially outperforming existing approaches.

biological networksgene geometryintervention-specific response

This study addresses the limited generalization of existing gene regulatory network inference methods under inductive settings and the absence of evaluation benchmarks aligned with biological experimental needs. The authors reformulate the task as a ranking-centric inductive graph completion problem and introduce the first co-evolutionary discrete diffusion framework that jointly models discrete gene expression states and regulatory interactions. To enhance training efficiency, they propose a TF-ALL subgraph sampling strategy. Furthermore, they establish BEELINE-KGC, the first inductive benchmark focused on discovering high-confidence novel regulatory relationships. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art models on BEELINE-KGC, and ablation studies confirm the effectiveness of each component.

gene regulatory networkinductive inferenceknowledge graph completion

Hot Scholars

SK

Smita Krishnaswamy

Yale University
Machine LearningData MiningManifold LearningDeep Learning
XJ

Xiang Ji

Princeton University
Machine LearningStatistical Learning TheoryReinforcement Learning
MR

Mariana Recamonde-Mendoza

Universidade Federal do Rio Grande do Sul/Hospital de Clínicas de Porto Alegre
Machine LearningData ScienceBioinformaticsComputational Biology
SK

Sun Kim

Professor, Seoul National University; CTO, AIGENDRUG Co. Ltd.
BioinformaticsMachine learningCloud systemsString pattern matching algorithms...
JL

Jie Liu

City University of Hong Kong
AI4HealthMLLMMedical Imaging