pathway enrichment analysis

Aggregating and statistically testing sets of genes or proteins to identify biological pathways associated with diseases or traits, and combining evidence (metadata, embeddings) to produce interpretable mechanistic attributions. It also covers evaluating learned gene–trait associations for robustness, statistical significance, and clinical relevance across cohorts.

pathwayenrichmentanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Accurately predicting gene–disease associations remains challenging due to sparse and heterogeneous biomedical knowledge. Method: This study systematically evaluates knowledge graph embedding (KGE) for this task, introducing the first unified five-step evaluation framework to comparatively assess link prediction versus supervised node-pair classification paradigms. It quantifies the impact of disease ontology semantic richness (e.g., DO, GO) and cross-ontology links on prediction performance. Contribution/Results: Link prediction consistently outperforms node-pair classification in capturing semantic structure and overall accuracy. Integrating cross-ontology links yields a substantial +4.2% improvement in prediction accuracy, whereas enriching disease ontology semantics alone provides only marginal gain (+0.8%). These findings establish link prediction as the superior paradigm for gene–disease association modeling and reveal the critical role of cross-ontology structural information—beyond isolated ontology semantics—in enhancing predictive performance.

Assesses impact of ontology enrichment on prediction accuracyCompares link prediction versus node-pair classification performanceEvaluates gene-disease link prediction methods using knowledge graphs

This work proposes a retrieval-augmented generation framework that integrates biomedical knowledge graphs with large language models to address the lack of interpretable, multi-step reasoning in analyzing gene interactions and their downstream pathway effects. The approach constructs a heterogeneous knowledge graph from sources such as KEGG and WikiPathways, then employs subgraph retrieval and structured prompting to guide the language model in evidence-driven multi-hop reasoning. By incorporating pathway perturbation propagation simulations, the framework enables end-to-end interpretable inference from gene–gene interactions to resultant pathway state changes. To our knowledge, this is the first application of retrieval-augmented generation in this domain, yielding consistent, evidence-backed, and mechanistically transparent predictions across diverse biological contexts.

biological networksgene interactioninterpretable reasoning

Traditional genome-wide association studies (GWAS) struggle to uncover causal disease mechanisms, while existing knowledge graph–enhanced GWAS (KGWAS) approaches rely on generic knowledge graphs that often introduce spurious associations. To address this limitation, this work proposes a novel KGWAS framework that integrates perturb-seq–derived, cell-type-specific gene interaction networks to construct context-specific knowledge graphs, replacing generic ones. By combining knowledge graph pruning, data-driven modeling of gene relationships, and rigorous statistical inference, the method significantly enhances the sparsity, biological coherence, and robustness of identified disease pathways—without compromising statistical power.

causal mechanismscell-type specificitydisease mechanism

Effect Size-Driven Pathway Meta-Analysis for Gene Expression Data

Jan 23, 2025
JA
Juan Antonio Villatoro-García
🏛️ University of Granada

Traditional gene-expression meta-analyses operate at the single-gene level, suffering from information loss and limited biological interpretability due to cross-platform gene absence and technical heterogeneity. To address this, we propose a pathway-level effect-size-driven meta-analysis paradigm. Our method introduces a novel aggregation strategy based on single-sample Gene Set Enrichment Analysis (ssGSEA), constructing pathway-level matrices that preserve both magnitude and directionality of effects. It integrates ssGSEA, effect-size-weighted meta-analysis, and R-based statistical modeling, implemented as an open-source CRAN R package. Validated across multiple datasets for systemic lupus erythematosus (SLE) and Parkinson’s disease, our approach significantly reduces false-positive rates while enhancing cross-platform comparability and biological interpretability of pathway activity—overcoming key limitations of conventional gene-centric meta-analysis.

Enabling pathway-level effect size analysis across multiple studiesIntegrating omics datasets with missing genes across platformsOvercoming limitations of individual gene-level meta-analysis approaches

Data-Driven Logistic Regression Ensembles With Applications in Genomics

Feb 17, 2021
AC
A. Christidis
🏛️ University of British Columbia | KU Leuven

Addressing the challenge of simultaneously achieving high prediction accuracy and biologically interpretable biomarker identification in high-dimensional genomic binary classification, this paper proposes the first data-driven logistic regression (LR) ensemble framework that jointly ensures statistical interpretability and strong predictive performance. The method integrates L₁/L₂ regularization with ensemble learning to directly learn a small set of highly accurate and interpretable LR models via global optimization. It provides the first rigorous derivation of asymptotic statistical properties for such regularized ensemble LR estimators. Additionally, we develop a gene importance ranking tool based on resampling stability to enhance biological interpretability. Evaluated on real-world datasets—including cancer, multiple sclerosis, and psoriasis—the framework achieves significant improvements in classification accuracy and successfully identifies several critical disease-associated genes—previously missed by competing methods—that are independently validated in the biomedical literature.

Develops data-driven logistic regression ensembles for genomicsImproves prediction accuracy and biomarker identification in diseasesProvides variable importance ranking for prioritizing critical genes

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently translating complex biomarker mechanisms—burdened by the exponential growth of biomedical literature and databases—into testable drug combination hypotheses. To this end, we propose CoDHy, a human–AI collaborative research system that integrates structured databases and unstructured literature to construct a task-oriented knowledge graph. CoDHy introduces a novel framework combining knowledge graph embeddings with agent-based reasoning to enable traceable, intervenable, and transparent generation, validation, and ranking of drug combination hypotheses. Through an interactive interface and an end-to-end workflow, CoDHy effectively supports researcher-driven exploratory hypothesis generation and decision-making in translational oncology, with its feasibility and practical utility demonstrated in real-world scenarios.

biomarkerbiomedical literaturedrug combination

In cancer genomics, low-statistical-power cohorts often yield false-negative driver genes and false-positive passenger genes due to stringent multiple-testing correction. To address this, we propose a multi-evidence integration framework that transcends conventional p-value–driven approaches by jointly leveraging causal inference techniques—specifically inverse probability weighting and doubly robust estimation—and multidimensional biological evidence, including mutational signatures, expression perturbations, and literature-supported functional annotations. We formalize these components into a five-criteria decision model. Our framework markedly improves robustness and specificity in detecting weak oncogenic signals. Applied to TCGA breast cancer data, it successfully identifies *KMT2C* as a candidate driver gene while excluding well-known false positives such as *RYR2*. The method is both interpretable—enabling transparent evidence tracing—and scalable to diverse genomic data types. It establishes a novel paradigm for driver event discovery in small-sample cancer cohorts.

Addresses false negatives and false positives in underpowered cohortsDistinguishes true biological drivers from statistical artifactsRescues low-power prognostic signals in cancer genomics studies

This study addresses the risk of introducing non-causal features and collider bias when constructing molecular signatures of lifestyle exposures without accounting for underlying causal structures. Leveraging directed acyclic graphs (DAGs) and d-separation theory, the work systematically elucidates, for the first time, the causal implications of univariate screening in feature construction and proposes incorporating this step prior to multivariable modeling to mitigate bias. Simulation studies demonstrate that while this strategy slightly reduces sensitivity and the correlation between exposure and signature, it substantially decreases the inclusion of non-causal features, yielding a feature set more aligned with the underlying causal mechanisms. Consequently, the approach offers clear advantages for mechanistic investigations seeking biologically interpretable signatures.

causalitycollider biaslifestyle exposures

Current methods for rare-variant association analysis exhibit limited statistical power when applied to survival outcomes and struggle to accommodate heterogeneous genetic effects. To address these limitations, this work proposes aSPU—a novel, data-adaptive aggregation testing framework tailored for survival data—built upon Schoenfeld residuals from the Cox model. The method flexibly integrates heterogeneity in both magnitude and direction of genetic effects at the gene and pathway levels. aSPU maintains high statistical power across diverse genetic architectures and is accompanied by an efficient R package enabling rapid computation and simulation. In an application to post-radiotherapy bladder toxicity in prostate cancer patients, the approach successfully replicated known genetic signals and identified novel biologically relevant genes and pathways.

aggregate association testsgenetic architecturesrare-variant associations

Traditional approaches to disease–gene association prediction rely heavily on manual literature curation, which is labor-intensive and poorly scalable. To address this limitation, this work proposes a graph neural network framework operating on a heterogeneous biological graph that integrates ProtT5 protein sequence embeddings with BioBERT disease text embeddings for the first time. The method employs a multi-relational graph learning architecture within an encoder–decoder paradigm to predict novel associations. Benchmark evaluations demonstrate that the proposed model significantly outperforms 14 state-of-the-art methods. Moreover, high-confidence predictions generated by the model are corroborated by existing literature and exhibit clear biological relevance, thereby offering a powerful tool for identifying candidate disease-causing genes.

biomedical knowledge graphdisease-gene associationgene-disease prediction

Hot Scholars

YZ

Yuanchun Zhou

Computer Network Information Center,CAS
Data MiningBig Data Analysis
ZY

Ziwei Yang

Bioinformatics Center, Institute for Chemical Research, Kyoto University
BioinformaticsMachine LearningComputational BiologyBiomedical Data Science
MW

Min Wu

Professor, IEEE Fellow, China University of Geosciences
Process controlRobust controlIntelligent systems
HW

Haohan Wang

School of Information Sciences, University of Illinois Urbana-Champaign
Computational BiologyAgentic AIAI4ScienceAI security
DR

David Richmond

AI and Machine Learning Scientist
computer vision for biomedical images