genomics data analysis

Empirically evaluating and applying statistical and computational methods to real and simulated genomics datasets—particularly sparse, high-dimensional data—to assess accuracy, efficiency, and to identify biologically relevant biomarkers.

genomicsdataanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

On the (In)Significance of Feature Selection in High-Dimensional Datasets

Aug 05, 2025
BN
Bhavesh Neekhra
🏛️ Ashoka University | Indian Institute of Technology Kharagpur

The necessity and efficacy of feature selection (FS) in high-dimensional gene expression classification remain widely assumed but insufficiently validated empirically. Many computational FS studies select genes without experimental verification, raising methodological concerns. Method: We propose a hypothesis-testing framework to systematically compare the classification performance of small, randomly sampled feature subsets (0.02%–1% of total features) against those selected by classical FS algorithms and the full feature set. Contribution/Results: On most benchmark gene expression datasets, classifiers trained on random feature subsets achieve accuracy comparable to—or even exceeding—that of FS-based and full-feature models. These findings challenge the prevailing assumption that FS inherently improves predictive performance. They expose potential methodological risks in computation-driven gene selection practices and provide empirical evidence urging critical reevaluation of FS conventions in biomedical feature engineering.

Challenges validity of gene selection in computational genomicsCompares random feature subsets to FS algorithm resultsEvaluates effectiveness of feature selection in high-dimensional datasets

Data-Driven Logistic Regression Ensembles With Applications in Genomics

Feb 17, 2021
AC
A. Christidis
🏛️ University of British Columbia | KU Leuven

Addressing the challenge of simultaneously achieving high prediction accuracy and biologically interpretable biomarker identification in high-dimensional genomic binary classification, this paper proposes the first data-driven logistic regression (LR) ensemble framework that jointly ensures statistical interpretability and strong predictive performance. The method integrates L₁/L₂ regularization with ensemble learning to directly learn a small set of highly accurate and interpretable LR models via global optimization. It provides the first rigorous derivation of asymptotic statistical properties for such regularized ensemble LR estimators. Additionally, we develop a gene importance ranking tool based on resampling stability to enhance biological interpretability. Evaluated on real-world datasets—including cancer, multiple sclerosis, and psoriasis—the framework achieves significant improvements in classification accuracy and successfully identifies several critical disease-associated genes—previously missed by competing methods—that are independently validated in the biomedical literature.

Develops data-driven logistic regression ensembles for genomicsImproves prediction accuracy and biomarker identification in diseasesProvides variable importance ranking for prioritizing critical genes

Automated Statistical and Machine Learning Platform for Biological Research

Nov 25, 2025
LR
Luke Rimmo Lego
🏛️ Stevens Institute of Technology

In biological research, the fragmentation between statistical analysis and machine learning tools, coupled with high usability barriers for non-programming users, impedes efficient and rigorous data-driven discovery. Method: We propose BioAutoML, a modular, biology-oriented automated analysis platform integrating classical statistical methods (e.g., t-tests, ANOVA, Pearson correlation) with interpretable machine learning (e.g., Random Forest classification). It supports automated data preprocessing, categorical encoding, feature importance assessment, and data-aware dynamic model configuration. Crucially, it introduces the first unified statistical–machine learning workflow, bridging methodological gaps via automated hyperparameter optimization. Contribution/Results: Evaluated on multiple chemomics datasets, BioAutoML achieves significantly higher classification accuracy than baseline approaches while preserving statistical validity. It enables domain scientists without programming expertise to perform end-to-end, interpretable, and statistically sound modeling—substantially accelerating biological insight generation.

Automates hyperparameter optimization and feature importance for non-programmersIntegrates statistical and machine learning methods for biological data analysisUnifies diverse tools to streamline workflows and enhance interpretability in bioinformatics

A robust and powerful replicability analysis for high dimensional data

Oct 15, 2023
HL
Haochen Lei
🏛️ Florida State University | Jilin University

Existing methods for assessing cross-study replicability of high-dimensional data suffer from strong parametric assumptions, poor scalability to multi-study settings, and insufficient statistical power. To address these limitations, this paper proposes the first empirical Bayes framework that simultaneously achieves high statistical power and strong robustness. Our method adaptively integrates information across multiple features and studies, explicitly models effect heterogeneity, and guarantees strict false discovery rate (FDR) control via theoretical analysis. Applied to real genome-wide association study (GWAS) data, it substantially improves detection of replicable signals and uncovers several biologically meaningful findings missed by mainstream approaches. Both theoretical analysis and empirical evaluation demonstrate that our method consistently outperforms state-of-the-art techniques in statistical power, robustness to model misspecification, and interpretability.

Addresses computational challenges in multi-study replicability with scalable pairwise strategy.Develops a method for replicability analysis across multiple high-dimensional studies.Identifies replicable signals in biomedical data, like genetic associations for diabetes.

Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation

May 06, 2025
DP
Diego Perazzolo
🏛️ University of Padova | I4 Consulting Srl | London South Bank University

To address poor interpretability and limited generalizability in few-shot, high-dimensional omics classification, this study proposes an interpretable classification framework integrating feature selection with synthetic data generation. Methodologically, it innovatively combines bootstrap-based feature screening, hybrid data augmentation using SMOTE and TABGAN, and an ensemble of binary classifiers, rigorously evaluated via stratified cross-validation. We systematically uncover, for the first time, the synergistic mechanism by which joint feature selection and synthetic data generation jointly enhance both interpretability and generalizability—even under extreme data scarcity (e.g., *n* < 30 per class). The framework demonstrates robust performance across six binary classification tasks in the EMTAB-8026 benchmark and maintains consistent accuracy when transferred to larger independent test sets, with no statistically significant degradation. This work establishes a novel paradigm for few-shot omics modeling that simultaneously ensures transparency, reliability, and practical utility.

Addressing limited samples in high-dimensional omics datasetsBalancing feature selection and synthetic data for reliabilityEnhancing omics classification accuracy and interpretability

Latest Papers

What's happening recently
View more

This study addresses the challenge of effectively integrating heterogeneous proteomic data—specifically, whole-sample mass spectrometry (MS) and multiplexed protein array profiles—in pancreatic cancer research. To overcome the limitations of conventional approaches that naively concatenate multi-source features, the authors propose a novel model fusion framework that explicitly models and leverages the heterogeneity between data sources through a tailored integration strategy. This approach synergistically exploits the complementary strengths of each modality rather than treating them as homogeneous inputs. Experimental results demonstrate that the proposed method significantly outperforms both single-modality models and standard fusion baselines in pancreatic cancer classification, yielding substantial improvements in diagnostic accuracy. The work thus offers a principled and effective paradigm for integrating heterogeneous multi-omics data in biomedical applications.

classification modelsdata integrationheterogeneous proteomic data

Clustering analyses often lack quantitative assessment of reproducibility. To address this gap, this work proposes ERICA, the first systematic framework for quantifying clustering reproducibility. ERICA generates stability statistics through iterative cluster assignments and integrates quantitative visualization to reveal inter-cluster similarity and potential outliers. The method is validated on synthetic datasets and applied to breast cancer gene expression data, where it identifies subsets of clustering results that are irreproducible. These findings underscore ERICA’s critical value in real-world applications for evaluating the reliability of clustering outcomes and the robustness of underlying data structures.

cluster analysisclusteringquantitative evaluation

To address the computational budget limitations in large-scale hypothesis testing—where exact $p$- or $e$-value computation is often infeasible—this paper proposes a budget-aware active testing framework. The method leverages auxiliary statistics to adaptively decide, via a probabilistic mechanism, whether to compute exact or efficient surrogate test statistics, ensuring strict budget adherence in expectation. We establish theoretical guarantees of statistical optimality and validity under both independence and dependence assumptions. The framework integrates principled $p$/$e$-value construction, randomized decision-making, and active sampling strategies. Extensive evaluations—including synthetic simulations, genome-wide association studies (GWAS), and large language model–driven clinical prediction tasks—demonstrate substantial gains in statistical power under fixed computational budgets, while preserving scalability and inferential reliability.

Develops budget-aware methods for large-scale hypothesis testing.Ensures valid p-values or e-values under budget constraints.Uses auxiliary statistics to allocate computational resources efficiently.

This study addresses the computational and interpretability challenges in joint modeling of high-dimensional predictors and responses, where existing approaches typically reduce dimensionality only on the predictor side. To overcome these limitations, this work proposes the Graph-Independent Dual Screening (GIDS) framework, which enables efficient simultaneous dimension reduction for both predictors and responses for the first time, accompanied by a theoretically guaranteed statistical screening algorithm. The GIDS method substantially enhances computational efficiency, model scalability, and result interpretability. Empirical evaluations demonstrate its superior performance over state-of-the-art methods in simulations. Applied to ADNI data, GIDS successfully reduces 860,000 CpG sites and 49,000 transcripts to approximately 9,000 and 2,000 features, respectively, revealing block-wise CpG–gene interaction structures associated with Alzheimer’s disease.

dimensionality reductionhigh-dimensional outcomeshigh-dimensional predictors

This study addresses the challenges of classifying breast cancer subtypes in high-dimensional gene expression data, where small sample sizes and class imbalance hinder performance. The authors systematically evaluate the impact of model complexity, feature dimensionality, and evaluation metrics on classification outcomes by applying logistic regression, random forest, and support vector machines across varying numbers of highly variable genes, using stratified five-fold cross-validation. Results demonstrate that feature dimensionality and choice of evaluation metric—particularly macro F1-score—are more influential than model complexity. Logistic regression exhibits the most stable and balanced performance across all subtypes, including rare ones, whereas random forest achieves higher overall accuracy but poorer minority-class recognition, and support vector machines prove highly sensitive to feature dimensionality. These findings offer practical guidance for model selection in high-dimensional biomedical subtype classification tasks.

breast cancer subtype classificationclass imbalancegene expression data

Hot Scholars

SA

Sarwan Ali

Columbia University
Deep LearningMachine LearningAdversarial AttackCombinatorial Optimization
CF

Can Firtina

Assistant Professor of Computer Science, UMD
BioinformaticsComputer ArchitectureHardware-Software Co-design
KK

Konstantina Koliogeorgi

National Technical University of Athens
FPGAsHigh level synthesisarchitectural optimizations
MP

Murray Patterson

Assistant Professor, Georgia State University
Computational BiologyAlgorithmsCombinatoricsData Science