Score
Design, implement, and evaluate computational methods that infer quantitative activity scores for biological pathways from high-dimensional molecular measurements (e.g., gene expression or multi‑omic data). This includes building pathway‑informed models—often autoencoders whose architecture or regularization is constrained by pathway membership—to produce interpretable pathway‑level latent features for downstream analyses such as classification, clustering, or outcome modeling.
This study addresses the challenge of balancing model expressiveness and interpretability in multi-omics data integration by proposing a Pathway Activity Autoencoder (PAA). The PAA embeds prior biological pathway knowledge directly into the network architecture as structural constraints, thereby achieving intrinsic interpretability without sacrificing predictive power. By integrating diverse omics data—including gene expression, protein abundance, and miRNA profiles—and leveraging pathway-guided architectural design together with tailored regularization strategies, the method significantly outperforms existing approaches in breast cancer survival prediction and molecular subtype classification. Experimental results demonstrate not only enhanced predictive performance but also clear attribution of each omics layer’s contribution to the predictions, offering robustness and clinically meaningful interpretability.
It remains unclear whether pathway-guided deep learning models improve performance due to biologically meaningful priors or merely because pathway structures induce beneficial sparsity. Method: We systematically constructed 15 pathway-based models and their rigorously matched random sparse variants—controlling for sparsity level, number of connections, and parameter initialization—and evaluated them across multiple datasets on both predictive performance and interpretability. Contribution/Results: Performance gains are predominantly attributable to sparse regularization rather than biological relevance: three random sparse models significantly outperformed their pathway-based counterparts while retaining accurate disease biomarker identification; pathway models showed no interpretability advantage. This work is the first to mechanistically disentangle the role of biological priors in pathway modeling, introducing a general randomized benchmarking framework. It provides a rigorous methodological foundation for attributing efficacy to domain-specific priors across scientific disciplines.
Traditional feature selection methods for high-dimensional genomic data neglect biological pathway structures, leading to unstable and biologically uninterpretable results. To address this, we propose a two-stage collaborative reinforcement learning framework that integrates statistical rigor with pathway prior knowledge. Our approach innovatively couples multi-agent reinforcement learning (MARL) with KEGG pathway annotations, employs graph neural networks (GNNs) to model gene–gene interactions, and designs a composite reward function balancing predictive accuracy and pathway coverage. We further introduce shared memory and a centralized critic to enable coordinated agent optimization. Evaluated across multiple gene expression datasets, our method achieves an average 4.2% improvement in disease classification accuracy, enhances pathway enrichment significance by 3.8×, and significantly improves cross-dataset stability and biological consistency.
Current pathway-guided models lack a unified benchmark for simultaneously predicting eligibility for targeted therapy, need for radiotherapy, and six-month survival. This study proposes the first integrative evaluation framework based on Reactome pathway activity scores, jointly training three bioinformatic architectures—BINN, GraphPath, and PATH—across five TCGA cancer cohorts to enable multitask clinical outcome prediction. It innovatively applies deep learning over pathway structures to jointly model therapeutic response and survival, while establishing a cross-model protocol for fair comparison. Results show that PATH achieves overall superior performance in targeted therapy prediction, BINN excels in survival prediction, and GraphPath attains an AUROC of 0.92 for targeted therapy prediction in prostate cancer with well-defined driver mutations. Radiotherapy prediction remains suboptimal, likely because key decision-making factors are not captured in gene expression data.
Traditional biomedical knowledge extraction relies heavily on manual curation, resulting in low scalability and efficiency. Method: This study conducts the first systematic evaluation of large language models (LLMs) for genome-scale molecular interaction and pathway knowledge extraction. We integrate BioBERT, LLaMA, and ChatGLM with prompt engineering and supervise fine-tuning using gold-standard databases—including STRING, KEGG, and Reactome—alongside zero-shot inference. Results: Large models significantly outperform smaller ones, achieving moderate performance (F1 ≈ 0.62) on protein–protein interaction identification, radiation-response pathway gene discovery, and gene regulatory relationship parsing. However, critical bottlenecks persist in identifying functionally heterogeneous gene clusters and modeling strongly correlated regulatory relationships. This work provides empirical evidence and methodological guidance for AI-driven, scalable, and automated biological knowledge discovery.
This work proposes LaCoGSEA, a novel framework that integrates deep autoencoders with gene set enrichment analysis (GSEA) to address the limitations of traditional pathway enrichment methods in unsupervised settings. Conventional approaches rely on predefined phenotypic labels and are thus ill-suited for label-free scenarios, while existing unsupervised methods often assume linearity and fail to explicitly model gene–pathway relationships. LaCoGSEA overcomes these issues by leveraging an autoencoder to capture the nonlinear manifold of transcriptomic data and generating label-free gene rankings based on global correlations between genes and latent variables. These rankings drive a GSEA-like enrichment statistic without requiring phenotype annotations. Evaluated on cancer subtype clustering tasks, LaCoGSEA significantly outperforms current unsupervised baselines, recovers more high-confidence biological pathways, and demonstrates robust performance across varying data scales and experimental conditions, establishing a new state of the art in unsupervised pathway enrichment analysis.
This work addresses the limitation of traditional tensor decomposition methods in incorporating prior knowledge from computational models when analyzing high-dimensional multi-way data, such as metabolomics datasets, which hinders the discovery of interpretable patterns. The authors propose a knowledge-guided coupled tensor decomposition framework that, for the first time, jointly analyzes real observational data and simulated data generated by computational models under linear coupling constraints. This approach enhances both robustness and interpretability of extracted patterns in noisy settings and successfully identifies latent inconsistencies between model predictions and empirical observations in real metabolomics data. The results demonstrate the method’s effectiveness and novelty in seamlessly integrating domain-specific prior knowledge with data-driven analysis.
This study addresses the risk of introducing non-causal features and collider bias when constructing molecular signatures of lifestyle exposures without accounting for underlying causal structures. Leveraging directed acyclic graphs (DAGs) and d-separation theory, the work systematically elucidates, for the first time, the causal implications of univariate screening in feature construction and proposes incorporating this step prior to multivariable modeling to mitigate bias. Simulation studies demonstrate that while this strategy slightly reduces sensitivity and the correlation between exposure and signature, it substantially decreases the inclusion of non-causal features, yielding a feature set more aligned with the underlying causal mechanisms. Consequently, the approach offers clear advantages for mechanistic investigations seeking biologically interpretable signatures.
Existing explainable AI methods for predicting cancer drug response suffer from high computational costs, limited robustness, and an exclusive focus on single-gene importance scores, which hinders the elucidation of synergistic gene network mechanisms. To address these limitations, this work proposes ILLUME+, a novel framework that, for the first time, integrates multiple complementary forms of explanation into an end-to-end pipeline. ILLUME+ combines transcriptome data–driven machine learning models with post-hoc interpretability algorithms and incorporates pathway analysis alongside mechanistic validation. The resulting approach yields more stable, multi-gene interaction–based explanations that not only recapitulate known drug–gene associations and underlying mechanisms but also uncover novel cooperative driver signals, thereby enhancing biological interpretability while maintaining scalability.
In biological research, the fragmentation between statistical analysis and machine learning tools, coupled with high usability barriers for non-programming users, impedes efficient and rigorous data-driven discovery. Method: We propose BioAutoML, a modular, biology-oriented automated analysis platform integrating classical statistical methods (e.g., t-tests, ANOVA, Pearson correlation) with interpretable machine learning (e.g., Random Forest classification). It supports automated data preprocessing, categorical encoding, feature importance assessment, and data-aware dynamic model configuration. Crucially, it introduces the first unified statistical–machine learning workflow, bridging methodological gaps via automated hyperparameter optimization. Contribution/Results: Evaluated on multiple chemomics datasets, BioAutoML achieves significantly higher classification accuracy than baseline approaches while preserving statistical validity. It enables domain scientists without programming expertise to perform end-to-end, interpretable, and statistically sound modeling—substantially accelerating biological insight generation.