batch effect correction

Preprocessing and statistical adjustment techniques to remove protocol- or batch-induced systematic variation across omics or similar datasets, enabling proper alignment, fusion, and downstream biological signal recovery.

batcheffectcorrection

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

MoDaH achieves rate optimal batch correction

Dec 09, 2025
YC
Yang Cao
🏛️ Yale University

In single-cell multi-omics analysis, batch effects introduce technical noise that obscures true biological signals, and existing batch correction methods lack rigorous statistical guarantees. To address this, we propose MoDaH—a batch normalization method grounded in an anisotropic Gaussian mixture model. MoDaH is the first to establish a minimax-optimal error rate lower bound for batch correction and provably achieves the information-theoretically optimal convergence rate. It integrates robust clustering theory with a certifiably convergent data harmonization optimization framework. On single-cell RNA-seq and spatial proteomics datasets, MoDaH matches or surpasses state-of-the-art methods—including Harmony, Seurat-v5, and LIGER—in correction accuracy, while strictly preserving biological heterogeneity and structural fidelity. This work bridges a critical gap in the statistical theory of batch effect correction.

Balances technical noise removal with biological signal preservationCorrects batch effects in single-cell omics dataProvides theoretical guarantees for batch correction reliability

This study addresses the lack of systematic evaluation in cross-modal integration of single-cell multi-omics data, where method performance is highly dependent on data characteristics. It presents the first comprehensive benchmark of combinations across seven normalization strategies, five integration methods—including Seurat and Harmony—and four dimensionality reduction techniques such as UMAP, spanning the entire pipeline from preprocessing to embedding. Using standardized metrics including Silhouette coefficient, Adjusted Rand Index (ARI), and Calinski–Harabasz index, the work reveals critical compatibilities and context-specific strengths: Harmony demonstrates superior computational efficiency on large-scale datasets, Seurat achieves higher integration accuracy, and UMAP exhibits the broadest compatibility across integration approaches. Importantly, the findings underscore that normalization strategies must be jointly selected with integration methods to optimize performance.

benchmarkingdata integrationmultimodal data

This study addresses the high sensitivity of feature selection in untargeted LC-MS metabolomics to preprocessing pipelines, which renders results vulnerable to analytical degrees of freedom. To mitigate this, we introduce multi-universe analysis—a novel, auditable, configuration-driven consensus framework. Building upon a ten-stage quality control filter, our approach integrates four distinct preprocessing strategies with four feature-ranking methods, coupled with bootstrap-based stability selection and label permutation testing to retain only features consistently identified across multiple analytical paths. Among 30,370 initial features, individual pipelines selected 4–20 features with low Jaccard consistency (as low as 0.05), whereas the multi-universe consensus yielded 15 robust features reproducible in at least two out of four pathways. Notably, one feature was stable across all pathways and showed no false positives in 50 permutation tests, substantially enhancing result reliability and reproducibility.

analytical degrees of freedomfeature selectionLC-MS metabolomics

Dimension Reduction with Locally Adjusted Graphs

Dec 19, 2024
YW
Yingfan Wang
🏛️ Duke University

In high-dimensional data dimensionality reduction, the initial similarity graph is unreliable due to the “curse of dimensionality” and information sparsity, impeding cluster separation—especially as dataset size increases. To address this, we propose LocalMAP, a novel algorithm that introduces a dynamic local subgraph extraction and online update mechanism. LocalMAP achieves fine-grained, adaptive refinement of the adjacency graph via embedding-driven subgraph sampling, local neighborhood-aware adaptive reweighting, and iterative graph optimization. Compared with conventional methods (e.g., t-SNE, UMAP), LocalMAP significantly improves clustering structure recovery accuracy. On large-scale transcriptomic datasets, it successfully disentangles biologically meaningful but previously confounded subpopulations, accurately identifying critical cell types that were either missed or erroneously merged in prior analyses. LocalMAP thus establishes a new paradigm for interpretable, scalable dimensionality reduction of high-dimensional biological data.

Addresses unreliable graph construction in high-dimensional data clusteringEnhances identification of overlooked clusters in large biological datasetsImproves cluster separation through dynamic local graph adjustments

TxPert: Leveraging Biochemical Relationships for Out-of-Distribution Transcriptomic Perturbation Prediction

May 20, 2025
FW
Frederik Wenkel
🏛️ Valence Labs | Recursion | University of British Columbia

This study addresses the challenge of improving out-of-distribution (OOD) generalization for predicting transcriptional responses to genetic perturbations—specifically, unseen single- and double-gene perturbations and novel cell lines. To overcome the limited experimental coverage that constrains existing methods, we introduce, for the first time, multi-source biological knowledge graphs to guide OOD modeling, establishing a rigorous benchmark framework that enforces strict cross-perturbation-type and cross-cell-line generalization. Our method integrates graph neural networks for knowledge graph encoding, multi-relational heterogeneous graph aggregation, OOD-aware training, and an interpretable attention mechanism. Across all three OOD settings, TxPert achieves a mean R² improvement of 12.7% over state-of-the-art baselines. We publicly release both the new benchmark and the model implementation.

Enhancing perturbation modeling evaluation standardsImproving out-of-distribution prediction using gene-gene knowledge graphsPredicting cellular responses to unseen genetic perturbations

Latest Papers

What's happening recently
View more

This study addresses the challenge of reverse-identifying molecular targets and compounds from single-cell transcriptional perturbation responses. The authors propose the first multi-task Transformer-based retrieval model tailored for single-cell perturbation data, which jointly learns target prediction and molecular embedding within a fixed compound library. The model performs end-to-end inverse inference using differential expression profiles relative to cell-type-specific DMSO controls and incorporates a structure–transcriptome alignment constraint to enhance representational consistency. Evaluated on the Tahoe-100M dataset, the model achieves a target Recall@10 of 0.408 and a compound Hit@1 of 0.129, significantly outperforming baseline methods and demonstrating its effectiveness in retrieving known perturbation pairs.

compound retrievalinverse problemperturbation signature

This study addresses the challenge of effectively integrating heterogeneous proteomic data—specifically, whole-sample mass spectrometry (MS) and multiplexed protein array profiles—in pancreatic cancer research. To overcome the limitations of conventional approaches that naively concatenate multi-source features, the authors propose a novel model fusion framework that explicitly models and leverages the heterogeneity between data sources through a tailored integration strategy. This approach synergistically exploits the complementary strengths of each modality rather than treating them as homogeneous inputs. Experimental results demonstrate that the proposed method significantly outperforms both single-modality models and standard fusion baselines in pancreatic cancer classification, yielding substantial improvements in diagnostic accuracy. The work thus offers a principled and effective paradigm for integrating heterogeneous multi-omics data in biomedical applications.

classification modelsdata integrationheterogeneous proteomic data

This study addresses the scarcity of structured, context-rich experimental data in targeted protein degradation (TPD), which has hindered the development of computational models. To overcome this limitation, the authors propose the first expert-in-the-loop large language model (LLM) agent framework tailored for TPD. By integrating lightweight prompt optimization, terminology-aware transfer, and a triangulation-based validation mechanism, the framework automatically extracts multidimensional information—including compounds, targets, recruiters, and critical experimental conditions—from scientific literature. Requiring only minimal annotated data, it achieves high-accuracy cross-task transfer. The resulting molecular glue and PROTAC databases are expanded by 81% and 92%, respectively, with expert-validated accuracy rates of 92% and 82.5%, substantially enhancing condition-aware modeling of degrader activity.

compound identifiersdatabase curationexperimental context

This study addresses the fragmentation of mass spectrometry data, annotations, and metadata in untargeted metabolomics, which hinders traceable and reusable knowledge generation. To overcome this limitation, the authors propose MetaboKG—a analysis-centric knowledge graph framework that integrates GNPS molecular networks, public repository metadata, and multiple ontologies (including MS, ChEBI, and NCBITaxon) into a unified semantic model grounded in PROV-O and SIO. The framework introduces an extended USI identifier to enable deferred binding and cross-analysis linking. Validated on 680 GNPS datasets, MetaboKG effectively supports complex queries related to biochemical enrichment, environment-specific patterns, and cross-instrument variability, thereby facilitating traceable annotation reuse and reproducible semantic exploration.

analytical integrationdata fragmentationknowledge graph

Existing single-cell foundation models struggle to accurately simulate gene perturbation effects and lack a unified framework for modeling multi-omics and spatial data. This work proposes a scalable, unified foundation model that integrates a multimodal Transformer architecture with LoRA fine-tuning and, for the first time, combines Metropolis-Hastings MCMC sampling with masked conditional distributions to enable continuous, interpretable simulation of transcriptional responses under perturbations—avoiding out-of-distribution artifacts caused by conventional token-based manipulations. The model supports 154 fine-grained cell-type annotations and incorporates spatial multi-omics data such as MERFISH and mass spectrometry imaging. It significantly outperforms current methods across multiple benchmarks—including PBMC68K and Replogle Perturb-seq—in tasks spanning cell annotation, perturbation response prediction, and multi-omics integration.

foundation modelin-silico perturbationout-of-distribution artifacts

Hot Scholars

YS

Yiren Song

PH.D student, National University of Singapore
Generative AIDiffusionUnified model
RC

Rita Cucchiara

Università degli Studi di Modena e Reggio Emilia, Italia
Computer VisionPattern RecognitionDeep LearningMultimedia
WD

Weisheng Dong

School of Artificial Intelligence, Xidian University, China
Image ProcessingComputer VisionDeep Learning
GL

Guandong Li

hfut
Hyspectral image,Computer vision,AIGC
YN

Ying Nian Wu

UCLA Department of Statistics and Data Science
Generative AIRepresentation learningComputer visionComputational neuroscience