Score
Empirically evaluating and applying statistical and computational methods to real and simulated genomics datasets—particularly sparse, high-dimensional data—to assess accuracy, efficiency, and to identify biologically relevant biomarkers.
Longitudinal omics data pose significant challenges for dynamic modeling and clinical translation due to their high dimensionality, temporal imbalance, and non-Gaussian distributional properties. To address these challenges, this study rigorously delineates the theoretical boundaries between time-series and longitudinal analysis, and establishes a methodology classification framework specifically tailored to omics characteristics—encompassing single-cell longitudinal modeling, multi-omics integration, network dynamics, and FDA-compliant analysis. We systematically unify linear and generalized linear mixed models, functional data analysis, Bayesian hierarchical modeling, survival analysis, and multi-view cross-platform fusion algorithms. The resulting methodological guide comprehensively addresses modeling assumptions, algorithmic suitability, and software implementation, delivering a reproducible, scalable analytical framework. This work substantially enhances the rigor, interpretability, and translational utility of complex longitudinal omics studies.
The necessity and efficacy of feature selection (FS) in high-dimensional gene expression classification remain widely assumed but insufficiently validated empirically. Many computational FS studies select genes without experimental verification, raising methodological concerns. Method: We propose a hypothesis-testing framework to systematically compare the classification performance of small, randomly sampled feature subsets (0.02%–1% of total features) against those selected by classical FS algorithms and the full feature set. Contribution/Results: On most benchmark gene expression datasets, classifiers trained on random feature subsets achieve accuracy comparable to—or even exceeding—that of FS-based and full-feature models. These findings challenge the prevailing assumption that FS inherently improves predictive performance. They expose potential methodological risks in computation-driven gene selection practices and provide empirical evidence urging critical reevaluation of FS conventions in biomedical feature engineering.
Addressing the challenge of simultaneously achieving high prediction accuracy and biologically interpretable biomarker identification in high-dimensional genomic binary classification, this paper proposes the first data-driven logistic regression (LR) ensemble framework that jointly ensures statistical interpretability and strong predictive performance. The method integrates L₁/L₂ regularization with ensemble learning to directly learn a small set of highly accurate and interpretable LR models via global optimization. It provides the first rigorous derivation of asymptotic statistical properties for such regularized ensemble LR estimators. Additionally, we develop a gene importance ranking tool based on resampling stability to enhance biological interpretability. Evaluated on real-world datasets—including cancer, multiple sclerosis, and psoriasis—the framework achieves significant improvements in classification accuracy and successfully identifies several critical disease-associated genes—previously missed by competing methods—that are independently validated in the biomedical literature.
In biological research, the fragmentation between statistical analysis and machine learning tools, coupled with high usability barriers for non-programming users, impedes efficient and rigorous data-driven discovery. Method: We propose BioAutoML, a modular, biology-oriented automated analysis platform integrating classical statistical methods (e.g., t-tests, ANOVA, Pearson correlation) with interpretable machine learning (e.g., Random Forest classification). It supports automated data preprocessing, categorical encoding, feature importance assessment, and data-aware dynamic model configuration. Crucially, it introduces the first unified statistical–machine learning workflow, bridging methodological gaps via automated hyperparameter optimization. Contribution/Results: Evaluated on multiple chemomics datasets, BioAutoML achieves significantly higher classification accuracy than baseline approaches while preserving statistical validity. It enables domain scientists without programming expertise to perform end-to-end, interpretable, and statistically sound modeling—substantially accelerating biological insight generation.
Existing methods for assessing cross-study replicability of high-dimensional data suffer from strong parametric assumptions, poor scalability to multi-study settings, and insufficient statistical power. To address these limitations, this paper proposes the first empirical Bayes framework that simultaneously achieves high statistical power and strong robustness. Our method adaptively integrates information across multiple features and studies, explicitly models effect heterogeneity, and guarantees strict false discovery rate (FDR) control via theoretical analysis. Applied to real genome-wide association study (GWAS) data, it substantially improves detection of replicable signals and uncovers several biologically meaningful findings missed by mainstream approaches. Both theoretical analysis and empirical evaluation demonstrate that our method consistently outperforms state-of-the-art techniques in statistical power, robustness to model misspecification, and interpretability.
To address poor interpretability and limited generalizability in few-shot, high-dimensional omics classification, this study proposes an interpretable classification framework integrating feature selection with synthetic data generation. Methodologically, it innovatively combines bootstrap-based feature screening, hybrid data augmentation using SMOTE and TABGAN, and an ensemble of binary classifiers, rigorously evaluated via stratified cross-validation. We systematically uncover, for the first time, the synergistic mechanism by which joint feature selection and synthetic data generation jointly enhance both interpretability and generalizability—even under extreme data scarcity (e.g., *n* < 30 per class). The framework demonstrates robust performance across six binary classification tasks in the EMTAB-8026 benchmark and maintains consistent accuracy when transferred to larger independent test sets, with no statistically significant degradation. This work establishes a novel paradigm for few-shot omics modeling that simultaneously ensures transparency, reliability, and practical utility.
This study addresses the challenge of effectively integrating heterogeneous proteomic data—specifically, whole-sample mass spectrometry (MS) and multiplexed protein array profiles—in pancreatic cancer research. To overcome the limitations of conventional approaches that naively concatenate multi-source features, the authors propose a novel model fusion framework that explicitly models and leverages the heterogeneity between data sources through a tailored integration strategy. This approach synergistically exploits the complementary strengths of each modality rather than treating them as homogeneous inputs. Experimental results demonstrate that the proposed method significantly outperforms both single-modality models and standard fusion baselines in pancreatic cancer classification, yielding substantial improvements in diagnostic accuracy. The work thus offers a principled and effective paradigm for integrating heterogeneous multi-omics data in biomedical applications.
Clustering analyses often lack quantitative assessment of reproducibility. To address this gap, this work proposes ERICA, the first systematic framework for quantifying clustering reproducibility. ERICA generates stability statistics through iterative cluster assignments and integrates quantitative visualization to reveal inter-cluster similarity and potential outliers. The method is validated on synthetic datasets and applied to breast cancer gene expression data, where it identifies subsets of clustering results that are irreproducible. These findings underscore ERICA’s critical value in real-world applications for evaluating the reliability of clustering outcomes and the robustness of underlying data structures.
To address the computational budget limitations in large-scale hypothesis testing—where exact $p$- or $e$-value computation is often infeasible—this paper proposes a budget-aware active testing framework. The method leverages auxiliary statistics to adaptively decide, via a probabilistic mechanism, whether to compute exact or efficient surrogate test statistics, ensuring strict budget adherence in expectation. We establish theoretical guarantees of statistical optimality and validity under both independence and dependence assumptions. The framework integrates principled $p$/$e$-value construction, randomized decision-making, and active sampling strategies. Extensive evaluations—including synthetic simulations, genome-wide association studies (GWAS), and large language model–driven clinical prediction tasks—demonstrate substantial gains in statistical power under fixed computational budgets, while preserving scalability and inferential reliability.
This study addresses the computational and interpretability challenges in joint modeling of high-dimensional predictors and responses, where existing approaches typically reduce dimensionality only on the predictor side. To overcome these limitations, this work proposes the Graph-Independent Dual Screening (GIDS) framework, which enables efficient simultaneous dimension reduction for both predictors and responses for the first time, accompanied by a theoretically guaranteed statistical screening algorithm. The GIDS method substantially enhances computational efficiency, model scalability, and result interpretability. Empirical evaluations demonstrate its superior performance over state-of-the-art methods in simulations. Applied to ADNI data, GIDS successfully reduces 860,000 CpG sites and 49,000 transcripts to approximately 9,000 and 2,000 features, respectively, revealing block-wise CpG–gene interaction structures associated with Alzheimer’s disease.
This study addresses the challenges of classifying breast cancer subtypes in high-dimensional gene expression data, where small sample sizes and class imbalance hinder performance. The authors systematically evaluate the impact of model complexity, feature dimensionality, and evaluation metrics on classification outcomes by applying logistic regression, random forest, and support vector machines across varying numbers of highly variable genes, using stratified five-fold cross-validation. Results demonstrate that feature dimensionality and choice of evaluation metric—particularly macro F1-score—are more influential than model complexity. Logistic regression exhibits the most stable and balanced performance across all subtypes, including rare ones, whereas random forest achieves higher overall accuracy but poorer minority-class recognition, and support vector machines prove highly sensitive to feature dimensionality. These findings offer practical guidance for model selection in high-dimensional biomedical subtype classification tasks.