Score
Preprocessing and statistical adjustment techniques to remove protocol- or batch-induced systematic variation across omics or similar datasets, enabling proper alignment, fusion, and downstream biological signal recovery.
In single-cell multi-omics analysis, batch effects introduce technical noise that obscures true biological signals, and existing batch correction methods lack rigorous statistical guarantees. To address this, we propose MoDaH—a batch normalization method grounded in an anisotropic Gaussian mixture model. MoDaH is the first to establish a minimax-optimal error rate lower bound for batch correction and provably achieves the information-theoretically optimal convergence rate. It integrates robust clustering theory with a certifiably convergent data harmonization optimization framework. On single-cell RNA-seq and spatial proteomics datasets, MoDaH matches or surpasses state-of-the-art methods—including Harmony, Seurat-v5, and LIGER—in correction accuracy, while strictly preserving biological heterogeneity and structural fidelity. This work bridges a critical gap in the statistical theory of batch effect correction.
This study addresses the lack of systematic evaluation in cross-modal integration of single-cell multi-omics data, where method performance is highly dependent on data characteristics. It presents the first comprehensive benchmark of combinations across seven normalization strategies, five integration methods—including Seurat and Harmony—and four dimensionality reduction techniques such as UMAP, spanning the entire pipeline from preprocessing to embedding. Using standardized metrics including Silhouette coefficient, Adjusted Rand Index (ARI), and Calinski–Harabasz index, the work reveals critical compatibilities and context-specific strengths: Harmony demonstrates superior computational efficiency on large-scale datasets, Seurat achieves higher integration accuracy, and UMAP exhibits the broadest compatibility across integration approaches. Importantly, the findings underscore that normalization strategies must be jointly selected with integration methods to optimize performance.
This study addresses the high sensitivity of feature selection in untargeted LC-MS metabolomics to preprocessing pipelines, which renders results vulnerable to analytical degrees of freedom. To mitigate this, we introduce multi-universe analysis—a novel, auditable, configuration-driven consensus framework. Building upon a ten-stage quality control filter, our approach integrates four distinct preprocessing strategies with four feature-ranking methods, coupled with bootstrap-based stability selection and label permutation testing to retain only features consistently identified across multiple analytical paths. Among 30,370 initial features, individual pipelines selected 4–20 features with low Jaccard consistency (as low as 0.05), whereas the multi-universe consensus yielded 15 robust features reproducible in at least two out of four pathways. Notably, one feature was stable across all pathways and showed no false positives in 50 permutation tests, substantially enhancing result reliability and reproducibility.
In high-dimensional data dimensionality reduction, the initial similarity graph is unreliable due to the “curse of dimensionality” and information sparsity, impeding cluster separation—especially as dataset size increases. To address this, we propose LocalMAP, a novel algorithm that introduces a dynamic local subgraph extraction and online update mechanism. LocalMAP achieves fine-grained, adaptive refinement of the adjacency graph via embedding-driven subgraph sampling, local neighborhood-aware adaptive reweighting, and iterative graph optimization. Compared with conventional methods (e.g., t-SNE, UMAP), LocalMAP significantly improves clustering structure recovery accuracy. On large-scale transcriptomic datasets, it successfully disentangles biologically meaningful but previously confounded subpopulations, accurately identifying critical cell types that were either missed or erroneously merged in prior analyses. LocalMAP thus establishes a new paradigm for interpretable, scalable dimensionality reduction of high-dimensional biological data.
This study addresses the challenge of improving out-of-distribution (OOD) generalization for predicting transcriptional responses to genetic perturbations—specifically, unseen single- and double-gene perturbations and novel cell lines. To overcome the limited experimental coverage that constrains existing methods, we introduce, for the first time, multi-source biological knowledge graphs to guide OOD modeling, establishing a rigorous benchmark framework that enforces strict cross-perturbation-type and cross-cell-line generalization. Our method integrates graph neural networks for knowledge graph encoding, multi-relational heterogeneous graph aggregation, OOD-aware training, and an interpretable attention mechanism. Across all three OOD settings, TxPert achieves a mean R² improvement of 12.7% over state-of-the-art baselines. We publicly release both the new benchmark and the model implementation.
This study addresses the challenge of reverse-identifying molecular targets and compounds from single-cell transcriptional perturbation responses. The authors propose the first multi-task Transformer-based retrieval model tailored for single-cell perturbation data, which jointly learns target prediction and molecular embedding within a fixed compound library. The model performs end-to-end inverse inference using differential expression profiles relative to cell-type-specific DMSO controls and incorporates a structure–transcriptome alignment constraint to enhance representational consistency. Evaluated on the Tahoe-100M dataset, the model achieves a target Recall@10 of 0.408 and a compound Hit@1 of 0.129, significantly outperforming baseline methods and demonstrating its effectiveness in retrieving known perturbation pairs.
This study addresses the challenge of effectively integrating heterogeneous proteomic data—specifically, whole-sample mass spectrometry (MS) and multiplexed protein array profiles—in pancreatic cancer research. To overcome the limitations of conventional approaches that naively concatenate multi-source features, the authors propose a novel model fusion framework that explicitly models and leverages the heterogeneity between data sources through a tailored integration strategy. This approach synergistically exploits the complementary strengths of each modality rather than treating them as homogeneous inputs. Experimental results demonstrate that the proposed method significantly outperforms both single-modality models and standard fusion baselines in pancreatic cancer classification, yielding substantial improvements in diagnostic accuracy. The work thus offers a principled and effective paradigm for integrating heterogeneous multi-omics data in biomedical applications.
This study addresses the scarcity of structured, context-rich experimental data in targeted protein degradation (TPD), which has hindered the development of computational models. To overcome this limitation, the authors propose the first expert-in-the-loop large language model (LLM) agent framework tailored for TPD. By integrating lightweight prompt optimization, terminology-aware transfer, and a triangulation-based validation mechanism, the framework automatically extracts multidimensional information—including compounds, targets, recruiters, and critical experimental conditions—from scientific literature. Requiring only minimal annotated data, it achieves high-accuracy cross-task transfer. The resulting molecular glue and PROTAC databases are expanded by 81% and 92%, respectively, with expert-validated accuracy rates of 92% and 82.5%, substantially enhancing condition-aware modeling of degrader activity.
This study addresses the fragmentation of mass spectrometry data, annotations, and metadata in untargeted metabolomics, which hinders traceable and reusable knowledge generation. To overcome this limitation, the authors propose MetaboKG—a analysis-centric knowledge graph framework that integrates GNPS molecular networks, public repository metadata, and multiple ontologies (including MS, ChEBI, and NCBITaxon) into a unified semantic model grounded in PROV-O and SIO. The framework introduces an extended USI identifier to enable deferred binding and cross-analysis linking. Validated on 680 GNPS datasets, MetaboKG effectively supports complex queries related to biochemical enrichment, environment-specific patterns, and cross-instrument variability, thereby facilitating traceable annotation reuse and reproducible semantic exploration.
Existing single-cell foundation models struggle to accurately simulate gene perturbation effects and lack a unified framework for modeling multi-omics and spatial data. This work proposes a scalable, unified foundation model that integrates a multimodal Transformer architecture with LoRA fine-tuning and, for the first time, combines Metropolis-Hastings MCMC sampling with masked conditional distributions to enable continuous, interpretable simulation of transcriptional responses under perturbations—avoiding out-of-distribution artifacts caused by conventional token-based manipulations. The model supports 154 fine-grained cell-type annotations and incorporates spatial multi-omics data such as MERFISH and mass spectrometry imaging. It significantly outperforms current methods across multiple benchmarks—including PBMC68K and Replogle Perturb-seq—in tasks spanning cell annotation, perturbation response prediction, and multi-omics integration.