Score
Computing quantitative scores that summarize the activity of biological pathways or gene sets from multi-sample molecular data to enable interpretation and comparison across cohorts. This entails encoding and preprocessing pathway-level signals and designing scoring or verifier functions that reflect biologically meaningful, clinically actionable effects at the single-cell or cohort level.
This study systematically evaluates the evolution, application efficacy, and methodological advances of generative AI (GenAI) in bioinformatics. Addressing six core questions—including “How does GenAI enhance accuracy and mechanistic interpretability in multi-omics analysis?”—it conducts a PRISMA-compliant literature review (2018–2024), integrating diffusion models, Transformer architectures, and heterogeneous biological data sources (e.g., UniProtKB, CELLxGENE). Results show that domain-specific pretrained models (e.g., ESM, AlphaFold3-derived architectures) significantly outperform general-purpose models in structural modeling, functional prediction, and synthetic data generation—particularly in context-aware representation learning and multimodal integration—thereby improving molecular representation fidelity. However, critical bottlenecks persist: limited scalability and pervasive data bias. The study proposes a novel paradigm—“biological-mechanism-driven generative modeling”—establishing a methodological foundation and technical roadmap for interpretable, verifiable GenAI–bioinformatics integration.
High patient heterogeneity in clinical trials impedes identification of biologically homogeneous subpopulations. To address this, we propose a semidefinite programming (SDP)-based subcohort discovery framework incorporating biologically interpretable constraints. Methodologically, we introduce the first SDP relaxation that integrates methylome–transcriptome co-variation priors and design a Goemans–Williamson-type randomized rounding algorithm with a theoretical approximation ratio of 0.82. Applied to the Curtis breast cancer cohort, our method identifies a clinically meaningful subcohort enriched for metastatic cases and systematically uncovers an interpretable subcohort characterized by coordinated hypermethylation of tumor suppressor genes and altered nuclear receptor expression. This work establishes a computationally rigorous, biologically grounded, and experimentally verifiable paradigm for mechanistic disease dissection and targeted therapeutic development.
Current cell state discovery relies on dimensionality reduction, visualization, and manual clustering interpretation; however, intra-cluster heterogeneity frequently compromises biomarker identification accuracy, resulting in high trial-and-error costs and poor interpretability. To address this, we propose a novel framework integrating Mixture-of-Experts (MoE) modeling with interactive visual analytics: the MoE model automatically learns nonlinear associations between cell subpopulations and gene biomarkers without imposing rigid clustering assumptions; concurrently, the visual interface enables biologists to iteratively formulate, test, and refine state hypotheses while incorporating domain knowledge to guide model optimization. Case studies on real single-cell datasets demonstrate that our approach significantly improves biomarker detection accuracy and biological interpretability, successfully aiding the discovery of novel cell states and reducing analytical uncertainty by 42% (per expert assessment) compared to conventional methods.
Current pathway-guided models lack a unified benchmark for simultaneously predicting eligibility for targeted therapy, need for radiotherapy, and six-month survival. This study proposes the first integrative evaluation framework based on Reactome pathway activity scores, jointly training three bioinformatic architectures—BINN, GraphPath, and PATH—across five TCGA cancer cohorts to enable multitask clinical outcome prediction. It innovatively applies deep learning over pathway structures to jointly model therapeutic response and survival, while establishing a cross-model protocol for fair comparison. Results show that PATH achieves overall superior performance in targeted therapy prediction, BINN excels in survival prediction, and GraphPath attains an AUROC of 0.92 for targeted therapy prediction in prostate cancer with well-defined driver mutations. Radiotherapy prediction remains suboptimal, likely because key decision-making factors are not captured in gene expression data.
This study addresses the challenge of improving molecular subtype characterization and clinical outcome prediction in breast cancer by jointly modeling protein sequence semantics and quantitative expression levels. We propose a novel integrative framework that fuses protein sequence embeddings—generated by ProtGPT2—with transcriptomic or proteomic expression data to construct biologically interpretable, discriminative multi-omics features. The method combines ensemble K-means clustering, XGBoost classification, PPI network analysis, and feature importance ranking to enable fine-grained subtype stratification and mechanistic insight extraction. It identifies key protein modules—including KMT2C, CLASP2, and MYO1B—that coordinately regulate hormonal signaling, cytoskeletal remodeling, and drug resistance pathways. In survival prediction and biomarker status classification tasks, our approach achieves F1-scores of 0.88 and 0.87, respectively—significantly outperforming conventional expression-only baselines.
Molecular profiling for cancer diagnosis and therapy selection typically relies on costly, invasive genomic assays. This study addresses the need for non-invasive, cost-effective alternatives using routine hematoxylin and eosin (H&E)-stained whole-slide images (WSIs). Method: We develop a multitask AI system built upon Virchow2—a foundation model pretrained on 3 million WSIs—and introduce pathological representation disentanglement coupled with clinical annotation alignment to enable pan-cancer molecular biomarker prediction from H&E slides alone. Contribution/Results: Our model simultaneously predicts 80 molecular biomarkers across diverse cancer types (mean AU-ROC = 0.89), encompassing alterations in 505 genes, activity of five core signaling pathways, DNA repair deficiency, tumor mutational burden (TMB), microsatellite instability (MSI), and chromosomal instability (CIN). Validated on 38,984 patients and 47,960 H&E slides, it identifies histological correlates for 40 biomarkers and links 58 to clinically actionable therapeutic targets—advancing digital pathology–driven precision oncology and companion diagnostics.
This work addresses the challenge of efficiently translating complex biomarker mechanisms—burdened by the exponential growth of biomedical literature and databases—into testable drug combination hypotheses. To this end, we propose CoDHy, a human–AI collaborative research system that integrates structured databases and unstructured literature to construct a task-oriented knowledge graph. CoDHy introduces a novel framework combining knowledge graph embeddings with agent-based reasoning to enable traceable, intervenable, and transparent generation, validation, and ranking of drug combination hypotheses. Through an interactive interface and an end-to-end workflow, CoDHy effectively supports researcher-driven exploratory hypothesis generation and decision-making in translational oncology, with its feasibility and practical utility demonstrated in real-world scenarios.
This study addresses the limited clinical translatability of AI models in computational pathology, which stems from insufficient interpretability and biological validation. To bridge this gap, the authors propose SAGE, an intelligent agent system that uniquely integrates literature-anchored reasoning with multi-agent collaboration to establish a closed-loop framework spanning hypothesis generation to empirical validation. By jointly modeling histopathology images, gene expression profiles, and clinical data through multimodal association, SAGE automatically discovers interpretable imaging biomarkers grounded in clear biological mechanisms and significantly associated with clinical outcomes. Experimental results demonstrate that SAGE substantially enhances the transparency, credibility, and clinical translatability of identified biomarkers, thereby advancing computational pathology models toward real-world clinical deployment.
This study addresses the challenge of balancing model expressiveness and interpretability in multi-omics data integration by proposing a Pathway Activity Autoencoder (PAA). The PAA embeds prior biological pathway knowledge directly into the network architecture as structural constraints, thereby achieving intrinsic interpretability without sacrificing predictive power. By integrating diverse omics data—including gene expression, protein abundance, and miRNA profiles—and leveraging pathway-guided architectural design together with tailored regularization strategies, the method significantly outperforms existing approaches in breast cancer survival prediction and molecular subtype classification. Experimental results demonstrate not only enhanced predictive performance but also clear attribution of each omics layer’s contribution to the predictions, offering robustness and clinically meaningful interpretability.
This work addresses the critical barriers to deploying clinical-grade AI biomarker models in computational pathology—namely, the absence of standardized intermediate representations, provenance tracking, and reproducible evaluation frameworks. To overcome these challenges, we establish the first shared benchmark framework for computational pathology, leveraging The Cancer Genome Atlas (TCGA) cohort to provide structured intermediate representations, predefined data splits, trained models, and evaluation metrics, with cross-validation on an independent Memorial Sloan Kettering Cancer Center (MSKCC) cohort. The framework integrates pathology foundation models (PFMs) to extract features from H&E whole-slide images, combined with multiple instance learning, quality control metadata, spatial coordinate mapping, and OncoKB annotations. Across 33 tumor–biomarker tasks, a high-performing subset of eight tasks achieved mean AUROCs of 0.831 on TCGA and 0.801 on MSKCC, demonstrating cross-institutional stability and establishing a reproducible, comparable foundation for AI-driven biomarker development.
This study addresses the challenge of constructing interpretable, robust biomarker-based decision rules in clinical practice that satisfy a prespecified positive predictive value (PPV) constraint. The authors propose a novel linear decision framework that maximizes the true positive rate (TPR) under a strict PPV guarantee while adaptively incorporating external individual risk information to enhance discriminative performance. To the best of our knowledge, this is the first method to achieve statistically optimal TPR under a PPV constraint, balancing clinical utility with theoretical rigor. Through constrained optimization modeling, an adaptive information fusion mechanism, asymptotic theoretical analysis, and finite-sample simulations, the approach demonstrates superior performance in numerical experiments and is successfully applied to develop an early screening rule for pancreatic ductal adenocarcinoma among newly diagnosed diabetic patients.