preprocess single-cell data

Designs and implements end-to-end preprocessing workflows that transform raw and near‑raw single-cell data into cleaned, normalized, and analysis-ready matrices by performing quality control, filtering, count normalization, batch correction, and feature selection. Builds informative proxy embeddings and packages datasets with metadata and auxiliary resources (such as cell and feature annotations) to support downstream analysis and reproducible reuse.

preprocesssingle-celldata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic evaluation in cross-modal integration of single-cell multi-omics data, where method performance is highly dependent on data characteristics. It presents the first comprehensive benchmark of combinations across seven normalization strategies, five integration methods—including Seurat and Harmony—and four dimensionality reduction techniques such as UMAP, spanning the entire pipeline from preprocessing to embedding. Using standardized metrics including Silhouette coefficient, Adjusted Rand Index (ARI), and Calinski–Harabasz index, the work reveals critical compatibilities and context-specific strengths: Harmony demonstrates superior computational efficiency on large-scale datasets, Seurat achieves higher integration accuracy, and UMAP exhibits the broadest compatibility across integration approaches. Importantly, the findings underscore that normalization strategies must be jointly selected with integration methods to optimize performance.

benchmarkingdata integrationmultimodal data

scUnified: An AI-Ready Standardized Resource for Single-Cell RNA Sequencing Analysis

Sep 30, 2025
PX
Ping Xu
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences | Columbia University

Current single-cell RNA sequencing (scRNA-seq) data lack standardized, ready-to-use resources; heterogeneous formats, inconsistent preprocessing pipelines, and divergent annotation strategies severely hinder reproducibility and fair benchmarking of computational methods. To address this, we introduce scUnified—a first-of-its-kind, cross-species (human/mouse), multi-tissue (nine tissues), AI-ready standardized scRNA-seq resource integrating 13 high-quality datasets. All data undergo uniform quality control, gene filtering, normalization, and batch correction, and are released in H5AD format—fully compatible with Scanpy, Seurat, and other mainstream analysis frameworks. Rigorous quality validation and empirical evaluation demonstrate that scUnified significantly improves stability and reproducibility in key tasks including cell clustering and marker gene identification. By providing a rigorously curated, harmonized benchmark dataset, scUnified establishes a reliable foundation for methodological benchmarking and fosters equitable, transparent evaluation of scRNA-seq analysis tools.

Data format variations hinder reproducibility and systematic comparisonsLack of uniform preprocessing complicates computational analysis workflowsStandardized datasets are scarce for scRNA-seq method evaluation

Single-cell pre-trained models often suffer from limited generalization due to the long-tailed distribution of cell types and covariate shifts in gene expression data. To address this, this work proposes CellRefine, a novel framework that introduces a prototype-guided post-pretraining phase between pretraining and fine-tuning. For the first time, CellRefine incorporates curated marker gene sets as biologically informed structural priors and refines the latent cell embedding manifold through multi-objective optimization. This approach substantially enhances the generalization performance of single-cell foundation models across diverse downstream tasks, achieving performance gains of up to 15%.

covariate shiftgeneralizationlong-tailed distribution

xML-workFlow: an end-to-end explainable scikit-learn workflow for rapid biomedical experimentation

Apr 02, 2025
KA
Khoa A. Tran
🏛️ QIMR Berghofer Medical Research Institute

Biomedical machine learning modeling often suffers from high computational resource consumption, poor code reusability, and insufficient reproducibility and traceability. To address these challenges, we propose an end-to-end, interpretable, lightweight ML workflow that innovatively integrates scikit-learn, MLflow, and SHAP—enabling automated experiment tracking, model training, post-hoc interpretability analysis, and modular extensibility. Designed as a template-based framework, it supports seamless cross-project transfer, significantly improving modeling efficiency, result reproducibility, and team collaboration. The workflow is open-sourced and has been adopted by multiple bioinformatics teams for disease prediction and multi-omics analysis tasks. It establishes the first standardized, production-ready ML engineering practice tailored to biomedical research, bridging a critical gap between methodological innovation and scalable, transparent, and maintainable ML deployment in the domain.

Enhances reproducibility traceability in biomedical ML projectsProvides rapid scalable ML workflow for biomedical researchReduces time effort in building iterating ML models

This study addresses the lack of interpretable, auditable, and domain-informed automated reasoning methods in single-cell RNA sequencing analysis. The authors propose an “omics-native reasoning” paradigm and develop the first framework enabling large language models to directly invoke single-cell data and bioinformatics tools within natural language dialogues. Key tasks—such as cell type annotation, developmental trajectory reconstruction, and transcription factor target inference—are reformulated as iterative, stepwise reasoning processes that support correction and refinement. By integrating multi-turn reasoning, dynamic tool invocation, and evaluation on the scBench benchmark, the approach ensures transparent and traceable analytical logic. Experiments demonstrate that iterative reasoning improves cell type annotation accuracy by 11% over one-shot prompting, reduces graph edit distance by 30% in trajectory reconstruction using Gemini-2.5-Pro, and effectively resolves ambiguities in marker gene interpretation and regulatory mechanisms.

cell-type annotationdevelopmental trajectoryomics-native reasoning

Latest Papers

What's happening recently
View more

This study addresses the critical yet underexamined role of data filtering in clinical machine learning, which alters statistical structures and directly impacts task complexity and model performance. Despite these effects, existing research frequently treats filtering as routine preprocessing with insufficient transparency. This work reconceptualizes data filtering as a core component of the scientific method, advocating its integration into the broader research paradigm rather than its treatment as a mere technical step. To this end, we develop a transparent and interpretable clinical data preprocessing pipeline and release the corresponding code as open source. Our analysis elucidates the mechanisms through which filtering decisions critically influence data distributions and downstream model efficacy. Ultimately, this research provides a novel framework for enhancing methodological rigor and reproducibility in clinical artificial intelligence studies.

Clinical Machine LearningData FilteringExplainability

This study investigates whether single-cell annotation methods can be misled by manipulating the composition of neighboring cells without altering the expression profile of target cells. To this end, we propose CohortHijack, a novel robustness auditing framework that reveals, for the first time, the query cohort composition as an attack surface that preserves target features. Our approach combines random and structured cell removal strategies with greedy, multi-start, and beam search algorithms to evaluate neighborhood- or clustering-based classifiers—specifically logistic regression and calibrated linear SVM—on the PBMC3K and Paul15 datasets. Experiments demonstrate that removing only a small fraction of non-target cells (average perturbation <0.4%) suffices to flip 19.67%–24.33% of target labels; this effect vanishes when neighborhood mechanisms are disabled, confirming the attack’s specific dependence on local cellular context.

adversarial manipulationcohort compositionneighborhood refinement

This work addresses the high storage and reuse costs of single-cell data and the lack of auditability in existing distillation methods that produce non-traceable synthetic data. The authors propose Minmax-CF, a method that, under fixed budgets of cells and genes, constructs a traceable core set by selecting only real observed cells through discrete minimax optimization and matching with static feature functions. This approach fully preserves original cell identifiers, gene symbols, and associated counts, labels, and metadata. Evaluated across multiple datasets, Minmax-CF achieves up to 96.52% balanced accuracy, substantially reduces pathway errors, and yields up to 2.55× GPU acceleration. The selected cells further enable efficient anomaly analysis and model validation.

auditabledata provenancereal-cell coresets

This study addresses the scarcity of structured, context-rich experimental data in targeted protein degradation (TPD), which has hindered the development of computational models. To overcome this limitation, the authors propose the first expert-in-the-loop large language model (LLM) agent framework tailored for TPD. By integrating lightweight prompt optimization, terminology-aware transfer, and a triangulation-based validation mechanism, the framework automatically extracts multidimensional information—including compounds, targets, recruiters, and critical experimental conditions—from scientific literature. Requiring only minimal annotated data, it achieves high-accuracy cross-task transfer. The resulting molecular glue and PROTAC databases are expanded by 81% and 92%, respectively, with expert-validated accuracy rates of 92% and 82.5%, substantially enhancing condition-aware modeling of degrader activity.

compound identifiersdatabase curationexperimental context

Hot Scholars

PZ

Peijie Zhou

Center for Machine Learning Research & Center for Quantitative Biology, Peking University
computational systems biologysingle-cell multiomicsstochastic modelsnonlinear dynamics
SK

Smita Krishnaswamy

Yale University
Machine LearningData MiningManifold LearningDeep Learning
ZZ

Zelin Zang

Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences
Deep Learning
YZ

Yuanchun Zhou

Computer Network Information Center,CAS
Data MiningBig Data Analysis
ZL

Zhen Lei

Associate Professor, OSCO Research Chair in Off-site Construction
Offsite ConstructionConstruction Engineering and Management