Score
Designs and implements methods and evaluation pipelines to discover, select, and validate biological or clinical biomarkers—individual molecular or derived features—that reliably predict outcomes. This includes applying wrapper and embedded feature‑selection algorithms, ranking candidates by predictive importance, producing compact interpretable biomarker subsets, and quantifying their predictive performance and association with outcomes.
This study addresses the instability of feature selection in biomedical data caused by high dimensionality, small sample sizes, multicollinearity, and missing values. To this end, we developed ROOFS, a Python toolkit that, for the first time, integrates multidimensional evaluation metrics—including realistic performance estimation based on semi-synthetic data—to systematically benchmark diverse feature selection methods. Our framework comprehensively assesses downstream predictive performance, stability, and feature reliability through variance inflation factor-based pre-screening, Benjamini–Hochberg–corrected joint statistical testing, optimism-corrected performance evaluation, and semi-synthetic data generation. Applied to the PIONeeR lung cancer immunotherapy dataset, ROOFS identified an optimal method among 253 models, significantly outperforming mainstream approaches such as LASSO and thereby enhancing the robustness and clinical translatability of biomarker discovery.
To address poor interpretability and limited generalizability in few-shot, high-dimensional omics classification, this study proposes an interpretable classification framework integrating feature selection with synthetic data generation. Methodologically, it innovatively combines bootstrap-based feature screening, hybrid data augmentation using SMOTE and TABGAN, and an ensemble of binary classifiers, rigorously evaluated via stratified cross-validation. We systematically uncover, for the first time, the synergistic mechanism by which joint feature selection and synthetic data generation jointly enhance both interpretability and generalizability—even under extreme data scarcity (e.g., *n* < 30 per class). The framework demonstrates robust performance across six binary classification tasks in the EMTAB-8026 benchmark and maintains consistent accuracy when transferred to larger independent test sets, with no statistically significant degradation. This work establishes a novel paradigm for few-shot omics modeling that simultaneously ensures transparency, reliability, and practical utility.
Molecular profiling for cancer diagnosis and therapy selection typically relies on costly, invasive genomic assays. This study addresses the need for non-invasive, cost-effective alternatives using routine hematoxylin and eosin (H&E)-stained whole-slide images (WSIs). Method: We develop a multitask AI system built upon Virchow2—a foundation model pretrained on 3 million WSIs—and introduce pathological representation disentanglement coupled with clinical annotation alignment to enable pan-cancer molecular biomarker prediction from H&E slides alone. Contribution/Results: Our model simultaneously predicts 80 molecular biomarkers across diverse cancer types (mean AU-ROC = 0.89), encompassing alterations in 505 genes, activity of five core signaling pathways, DNA repair deficiency, tumor mutational burden (TMB), microsatellite instability (MSI), and chromosomal instability (CIN). Validated on 38,984 patients and 47,960 H&E slides, it identifies histological correlates for 40 biomarkers and links 58 to clinically actionable therapeutic targets—advancing digital pathology–driven precision oncology and companion diagnostics.
For diseases such as hepatocellular carcinoma—where no single ideal biomarker exists—existing diagnostic models suffer from limited performance under skewed biomarker distributions, small inter-group differences, or insufficient sample sizes. To address this, we propose a parametric likelihood ratio–based multimarker diagnostic model: (i) it explicitly models the likelihood ratio as an interpretable diagnostic accuracy metric (e.g., sensitivity, specificity); (ii) it enables robust inference under missing data; and (iii) it incorporates a resource-aware biomarker selection mechanism. The method integrates multivariate statistical modeling, likelihood ratio optimization, and diagnostic evaluation (AUC, ROC analysis). Validated via extensive simulations and real clinical datasets, it significantly outperforms state-of-the-art classification and discriminant methods—particularly in low-sample-size and incomplete-data settings. An open-source R package ensures full reproducibility of results.
Addressing the challenge of simultaneously achieving high prediction accuracy and biologically interpretable biomarker identification in high-dimensional genomic binary classification, this paper proposes the first data-driven logistic regression (LR) ensemble framework that jointly ensures statistical interpretability and strong predictive performance. The method integrates L₁/L₂ regularization with ensemble learning to directly learn a small set of highly accurate and interpretable LR models via global optimization. It provides the first rigorous derivation of asymptotic statistical properties for such regularized ensemble LR estimators. Additionally, we develop a gene importance ranking tool based on resampling stability to enhance biological interpretability. Evaluated on real-world datasets—including cancer, multiple sclerosis, and psoriasis—the framework achieves significant improvements in classification accuracy and successfully identifies several critical disease-associated genes—previously missed by competing methods—that are independently validated in the biomedical literature.
This study addresses the challenge of constructing interpretable, robust biomarker-based decision rules in clinical practice that satisfy a prespecified positive predictive value (PPV) constraint. The authors propose a novel linear decision framework that maximizes the true positive rate (TPR) under a strict PPV guarantee while adaptively incorporating external individual risk information to enhance discriminative performance. To the best of our knowledge, this is the first method to achieve statistically optimal TPR under a PPV constraint, balancing clinical utility with theoretical rigor. Through constrained optimization modeling, an adaptive information fusion mechanism, asymptotic theoretical analysis, and finite-sample simulations, the approach demonstrates superior performance in numerical experiments and is successfully applied to develop an early screening rule for pancreatic ductal adenocarcinoma among newly diagnosed diabetic patients.
This study addresses the challenge of predicting radiotherapy sensitivity in non-small cell lung cancer (NSCLC) by establishing, for the first time, an integrative transcriptomic (RNA-seq) and proteomic (DIA-MS) analytical framework using SF2—the surviving fraction after 2 Gy irradiation—as the phenotypic endpoint. Leveraging Lasso-based feature selection coupled with support vector regression (SVR), the model was optimized via ten repetitions of five-fold cross-validation to enhance robustness. The integrative model achieved stable predictive performance across both omics layers (R² = 0.461–0.604), outperforming unimodal models. It identified 20 consistently dysregulated cross-omics biomarker genes enriched in DNA damage repair and cellular stress response pathways. This work not only validates the complementary value of multi-omics integration for mechanistic insight and clinical translation but also establishes a generalizable paradigm for radiobiological sensitivity prediction.
This study addresses the challenge of identifying compact and effective prognostic biomarkers from high-dimensional multi-omics data under limited sample sizes. To this end, the authors propose Sweeping*, an algorithm that innovatively employs a multi-view alternating optimization framework. By iteratively performing multi-objective optimization—guided by the concordance index and root sparsity—between single-view and multi-view representations, Sweeping* simultaneously selects key omics features and models cross-omics interactions, thereby avoiding information loss inherent in naive concatenation strategies. Leveraging the NSGA-III-CHS genetic algorithm, five-fold cross-validation, and Pareto front analysis, Sweeping* demonstrates superior trade-offs between predictive accuracy and model complexity compared to clinical baselines across three TCGA cohorts, validating both the efficacy of multi-omics integration and its cohort-specific dependencies.
This study addresses the challenge of effectively integrating heterogeneous proteomic data—specifically, whole-sample mass spectrometry (MS) and multiplexed protein array profiles—in pancreatic cancer research. To overcome the limitations of conventional approaches that naively concatenate multi-source features, the authors propose a novel model fusion framework that explicitly models and leverages the heterogeneity between data sources through a tailored integration strategy. This approach synergistically exploits the complementary strengths of each modality rather than treating them as homogeneous inputs. Experimental results demonstrate that the proposed method significantly outperforms both single-modality models and standard fusion baselines in pancreatic cancer classification, yielding substantial improvements in diagnostic accuracy. The work thus offers a principled and effective paradigm for integrating heterogeneous multi-omics data in biomedical applications.