Score
Aggregating and statistically testing sets of genes or proteins to identify biological pathways associated with diseases or traits, and combining evidence (metadata, embeddings) to produce interpretable mechanistic attributions. It also covers evaluating learned gene–trait associations for robustness, statistical significance, and clinical relevance across cohorts.
Accurately predicting gene–disease associations remains challenging due to sparse and heterogeneous biomedical knowledge. Method: This study systematically evaluates knowledge graph embedding (KGE) for this task, introducing the first unified five-step evaluation framework to comparatively assess link prediction versus supervised node-pair classification paradigms. It quantifies the impact of disease ontology semantic richness (e.g., DO, GO) and cross-ontology links on prediction performance. Contribution/Results: Link prediction consistently outperforms node-pair classification in capturing semantic structure and overall accuracy. Integrating cross-ontology links yields a substantial +4.2% improvement in prediction accuracy, whereas enriching disease ontology semantics alone provides only marginal gain (+0.8%). These findings establish link prediction as the superior paradigm for gene–disease association modeling and reveal the critical role of cross-ontology structural information—beyond isolated ontology semantics—in enhancing predictive performance.
This work proposes a retrieval-augmented generation framework that integrates biomedical knowledge graphs with large language models to address the lack of interpretable, multi-step reasoning in analyzing gene interactions and their downstream pathway effects. The approach constructs a heterogeneous knowledge graph from sources such as KEGG and WikiPathways, then employs subgraph retrieval and structured prompting to guide the language model in evidence-driven multi-hop reasoning. By incorporating pathway perturbation propagation simulations, the framework enables end-to-end interpretable inference from gene–gene interactions to resultant pathway state changes. To our knowledge, this is the first application of retrieval-augmented generation in this domain, yielding consistent, evidence-backed, and mechanistically transparent predictions across diverse biological contexts.
Traditional genome-wide association studies (GWAS) struggle to uncover causal disease mechanisms, while existing knowledge graph–enhanced GWAS (KGWAS) approaches rely on generic knowledge graphs that often introduce spurious associations. To address this limitation, this work proposes a novel KGWAS framework that integrates perturb-seq–derived, cell-type-specific gene interaction networks to construct context-specific knowledge graphs, replacing generic ones. By combining knowledge graph pruning, data-driven modeling of gene relationships, and rigorous statistical inference, the method significantly enhances the sparsity, biological coherence, and robustness of identified disease pathways—without compromising statistical power.
Traditional gene-expression meta-analyses operate at the single-gene level, suffering from information loss and limited biological interpretability due to cross-platform gene absence and technical heterogeneity. To address this, we propose a pathway-level effect-size-driven meta-analysis paradigm. Our method introduces a novel aggregation strategy based on single-sample Gene Set Enrichment Analysis (ssGSEA), constructing pathway-level matrices that preserve both magnitude and directionality of effects. It integrates ssGSEA, effect-size-weighted meta-analysis, and R-based statistical modeling, implemented as an open-source CRAN R package. Validated across multiple datasets for systemic lupus erythematosus (SLE) and Parkinson’s disease, our approach significantly reduces false-positive rates while enhancing cross-platform comparability and biological interpretability of pathway activity—overcoming key limitations of conventional gene-centric meta-analysis.
Addressing the challenge of simultaneously achieving high prediction accuracy and biologically interpretable biomarker identification in high-dimensional genomic binary classification, this paper proposes the first data-driven logistic regression (LR) ensemble framework that jointly ensures statistical interpretability and strong predictive performance. The method integrates L₁/L₂ regularization with ensemble learning to directly learn a small set of highly accurate and interpretable LR models via global optimization. It provides the first rigorous derivation of asymptotic statistical properties for such regularized ensemble LR estimators. Additionally, we develop a gene importance ranking tool based on resampling stability to enhance biological interpretability. Evaluated on real-world datasets—including cancer, multiple sclerosis, and psoriasis—the framework achieves significant improvements in classification accuracy and successfully identifies several critical disease-associated genes—previously missed by competing methods—that are independently validated in the biomedical literature.
This work addresses the challenge of efficiently translating complex biomarker mechanisms—burdened by the exponential growth of biomedical literature and databases—into testable drug combination hypotheses. To this end, we propose CoDHy, a human–AI collaborative research system that integrates structured databases and unstructured literature to construct a task-oriented knowledge graph. CoDHy introduces a novel framework combining knowledge graph embeddings with agent-based reasoning to enable traceable, intervenable, and transparent generation, validation, and ranking of drug combination hypotheses. Through an interactive interface and an end-to-end workflow, CoDHy effectively supports researcher-driven exploratory hypothesis generation and decision-making in translational oncology, with its feasibility and practical utility demonstrated in real-world scenarios.
In cancer genomics, low-statistical-power cohorts often yield false-negative driver genes and false-positive passenger genes due to stringent multiple-testing correction. To address this, we propose a multi-evidence integration framework that transcends conventional p-value–driven approaches by jointly leveraging causal inference techniques—specifically inverse probability weighting and doubly robust estimation—and multidimensional biological evidence, including mutational signatures, expression perturbations, and literature-supported functional annotations. We formalize these components into a five-criteria decision model. Our framework markedly improves robustness and specificity in detecting weak oncogenic signals. Applied to TCGA breast cancer data, it successfully identifies *KMT2C* as a candidate driver gene while excluding well-known false positives such as *RYR2*. The method is both interpretable—enabling transparent evidence tracing—and scalable to diverse genomic data types. It establishes a novel paradigm for driver event discovery in small-sample cancer cohorts.
This study addresses the risk of introducing non-causal features and collider bias when constructing molecular signatures of lifestyle exposures without accounting for underlying causal structures. Leveraging directed acyclic graphs (DAGs) and d-separation theory, the work systematically elucidates, for the first time, the causal implications of univariate screening in feature construction and proposes incorporating this step prior to multivariable modeling to mitigate bias. Simulation studies demonstrate that while this strategy slightly reduces sensitivity and the correlation between exposure and signature, it substantially decreases the inclusion of non-causal features, yielding a feature set more aligned with the underlying causal mechanisms. Consequently, the approach offers clear advantages for mechanistic investigations seeking biologically interpretable signatures.
Current methods for rare-variant association analysis exhibit limited statistical power when applied to survival outcomes and struggle to accommodate heterogeneous genetic effects. To address these limitations, this work proposes aSPU—a novel, data-adaptive aggregation testing framework tailored for survival data—built upon Schoenfeld residuals from the Cox model. The method flexibly integrates heterogeneity in both magnitude and direction of genetic effects at the gene and pathway levels. aSPU maintains high statistical power across diverse genetic architectures and is accompanied by an efficient R package enabling rapid computation and simulation. In an application to post-radiotherapy bladder toxicity in prostate cancer patients, the approach successfully replicated known genetic signals and identified novel biologically relevant genes and pathways.
Traditional approaches to disease–gene association prediction rely heavily on manual literature curation, which is labor-intensive and poorly scalable. To address this limitation, this work proposes a graph neural network framework operating on a heterogeneous biological graph that integrates ProtT5 protein sequence embeddings with BioBERT disease text embeddings for the first time. The method employs a multi-relational graph learning architecture within an encoder–decoder paradigm to predict novel associations. Benchmark evaluations demonstrate that the proposed model significantly outperforms 14 state-of-the-art methods. Moreover, high-confidence predictions generated by the model are corroborated by existing literature and exhibit clear biological relevance, thereby offering a powerful tool for identifying candidate disease-causing genes.