analyze single-cell data

Designs and applies end-to-end computational workflows to process and interpret single-cell datasets, performing preprocessing (quality control, filtering, normalization), feature selection, dimensionality reduction, clustering, cell-type annotation, trajectory/inference and differential expression, and dataset integration. Implements and evaluates methods and visualizations that extract and validate cell-resolved biological or phenotypic signals from single-cell sequencing or imaging-derived measurements.

analyzesingle-celldata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$184K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic evaluation in cross-modal integration of single-cell multi-omics data, where method performance is highly dependent on data characteristics. It presents the first comprehensive benchmark of combinations across seven normalization strategies, five integration methods—including Seurat and Harmony—and four dimensionality reduction techniques such as UMAP, spanning the entire pipeline from preprocessing to embedding. Using standardized metrics including Silhouette coefficient, Adjusted Rand Index (ARI), and Calinski–Harabasz index, the work reveals critical compatibilities and context-specific strengths: Harmony demonstrates superior computational efficiency on large-scale datasets, Seurat achieves higher integration accuracy, and UMAP exhibits the broadest compatibility across integration approaches. Importantly, the findings underscore that normalization strategies must be jointly selected with integration methods to optimize performance.

benchmarkingdata integrationmultimodal data

CellScout: Visual Analytics for Mining Biomarkers in Cell State Discovery.

Nov 24, 2025
RS
Rui Sheng
🏛️ Hong Kong University of Science and Technology | Westlake University | CAIR, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences | Zhejiang University | Singapore Management University

Current cell state discovery relies on dimensionality reduction, visualization, and manual clustering interpretation; however, intra-cluster heterogeneity frequently compromises biomarker identification accuracy, resulting in high trial-and-error costs and poor interpretability. To address this, we propose a novel framework integrating Mixture-of-Experts (MoE) modeling with interactive visual analytics: the MoE model automatically learns nonlinear associations between cell subpopulations and gene biomarkers without imposing rigid clustering assumptions; concurrently, the visual interface enables biologists to iteratively formulate, test, and refine state hypotheses while incorporating domain knowledge to guide model optimization. Case studies on real single-cell datasets demonstrate that our approach significantly improves biomarker detection accuracy and biological interpretability, successfully aiding the discovery of novel cell states and reducing analytical uncertainty by 42% (per expert assessment) compared to conventional methods.

Addresses inconsistencies in visual clustering of cellsIdentifies biomarkers for distinct cell statesUncovers hidden associations between cell populations and biomarkers

Single-cell RNA sequencing data exhibit diverse geometric structures—such as clusters, trajectories, and branches—yet existing analytical methods often assume predefined shapes and lack the capacity to automatically infer the intrinsic geometry of the data. To address this limitation, this work introduces scShapeBench, the first comprehensive benchmark specifically designed for geometric structure identification in single-cell data, comprising both synthetic and expert-annotated real datasets. Furthermore, the authors propose scReebTower, a novel method grounded in diffusion geometry that automatically detects data shape by extracting Reeb graphs. By integrating diffusion geometry with topological skeleton sampling, scReebTower bridges the gap between visualization and the selection of downstream analysis pipelines. Experimental results demonstrate that scReebTower outperforms existing approaches such as PAGA and Mapper across multiple datasets, confirming its effectiveness in automated geometric structure recognition.

automated analysishigh-dimensional datashape detection

cp_measure: API-first feature extraction for image-based profiling workflows

Jul 01, 2025
AF
Alán F. Muñoz
🏛️ Broad Institute of MIT and Harvard | Institute of Computational Biology, Helmholtz Zentrum München

Current bioimage analysis tools (e.g., CellProfiler) face bottlenecks in automated and reproducible feature extraction, hindering scalable deployment of machine learning workflows. To address this, we introduce *cp_measure*—a modular, API-first Python library that decouples CellProfiler’s core measurement engine and refactors it into a programmable interface, enabling seamless integration with the scientific Python ecosystem. The library supports end-to-end phenotypic analysis of 2D/3D cellular imaging and spatial transcriptomics data, ensuring highly consistent (Pearson correlation >0.999 with CellProfiler) and fully reproducible feature extraction. Empirical evaluation demonstrates its efficiency and scalability across multi-batch, multimodal biological image datasets. *cp_measure* significantly enhances robustness and productivity in feature-driven computational biology modeling, while preserving compatibility with established CellProfiler pipelines.

Automating image-based profiling for cellular state analysisIntegrating CellProfiler measurements with machine learning pipelinesOvercoming barriers in reproducible feature extraction workflows

This study addresses the lack of interpretable, auditable, and domain-informed automated reasoning methods in single-cell RNA sequencing analysis. The authors propose an “omics-native reasoning” paradigm and develop the first framework enabling large language models to directly invoke single-cell data and bioinformatics tools within natural language dialogues. Key tasks—such as cell type annotation, developmental trajectory reconstruction, and transcription factor target inference—are reformulated as iterative, stepwise reasoning processes that support correction and refinement. By integrating multi-turn reasoning, dynamic tool invocation, and evaluation on the scBench benchmark, the approach ensures transparent and traceable analytical logic. Experiments demonstrate that iterative reasoning improves cell type annotation accuracy by 11% over one-shot prompting, reduces graph edit distance by 30% in trajectory reconstruction using Gemini-2.5-Pro, and effectively resolves ambiguities in marker gene interpretation and regulatory mechanisms.

cell-type annotationdevelopmental trajectoryomics-native reasoning

Latest Papers

What's happening recently
View more

Current evaluations of bioinformatics agents overemphasize answer correctness while neglecting workflow auditability and scientific credibility. This work proposes a Function–Evidence–Validation (FEV) tri-dimensional evaluation framework centered on inspectable workflow trajectories, shifting the primary focus to workflow correctness for the first time. Through systematic literature review, trajectory analysis, and cross-domain benchmark mapping, the study comprehensively analyzes 109 agent systems and 28 evaluation resources across subfields including genomics, single-cell and spatial omics, and protein science. The findings reveal that while agents perform adequately in planning and execution, they exhibit significant deficiencies in reproducibility, traceability, external validation, and prospective experimental design. This research provides both theoretical grounding and practical guidance for developing transparent, auditable next-generation bioinformatics agents.

agentic bioinformaticsreproducibilityscientific credibility

This work addresses the pervasive challenges in bioinformatics tooling—such as fragmentation, complex dependencies, inconsistent documentation, and irreproducible environments—that severely hinder method reuse and adaptation. To overcome these limitations, the authors propose PoSyMed, an open modular platform that integrates biomedical workflows through formalized tool descriptions, containerized execution, a persistent workflow engine, and a conversational interface. Innovatively, a large language model is incorporated as a semantic assistant within a typed, validated, and human-supervised framework to support tool discovery, pipeline construction, and parameter configuration. This design significantly enhances analytical transparency and reproducibility. The platform’s efficacy is demonstrated in representative biomedical use cases, and it has been released as open-source software.

bioinformatics toolsexecution environmentreproducibility

This study addresses how spatial biologists can guide and validate complex tissue data analysis tasks executed by AI agents. Building upon the Claude Science agent, the authors employ contextual inquiry, formative pilots, and observational experiments to propose four key design directions: execution control, familiar views, source information transparency, and cross-environment accessible verification. The work reveals the epistemic mechanisms through which scientists rely on visual evidence to evaluate AI-generated results. Furthermore, it constructs a comprehensive empirical model of the analytical workflow encompassing both interactive control and verification. Ultimately, this research contributes a systematic design framework for human-AI collaborative scientific discovery, offering actionable insights into integrating intelligent agents within rigorous biological research practices.

Agentic WorkflowsAI-Assisted AnalysisHuman-AI Interaction

This work addresses the challenge of efficiently exploring and interpreting the high-dimensional combinatorial space of gene perturbations generated by AI-based virtual cell models and their complex transcriptional responses across diverse cell types. To this end, we propose a visual analytics system that, for the first time, integrates clustered overviews, compact glyph-based encodings, and coordinated multi-view interactions to enable systematic comparison and interpretable exploration of perturbation strategies. By incorporating AI-generated predictions and validating through real-world case studies and expert interviews, we demonstrate that our approach substantially enhances researchers’ understanding of perturbation effects and improves decision-making efficiency in drug discovery, effectively bridging the cognitive gap between computational models and biomedical experts.

drug discoverygene perturbationhigh-dimensional data

Hot Scholars

CC

Changxi Chi

Westlake University; NUAA
Deep LearningGenerative Model
SG

Stephan Günnemann

Professor of Computer Science, Technical University of Munich
Machine LearningGraphsGraph Neural NetworksRobustness
TC

Tianyu Cui

Research Scientist, Johnson and Johnson
Probabilistic ModelingDeep LearningDrug Discovery
JJ

Joakim Jaldén

Professor, KTH Royal Institute of Technology
Signal ProcessingCommunicationsInformation TheoryMachine Learning
LC

Long Cai

Research Professor of Biology, Caltech
single cell genomicsspatial genomics