TCellAlign: Cross-study T-cell Populations Alignment with Nomenclature-Guided Multi-Agent Workflow

๐Ÿ“… 2026-07-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of cross-study comparability in single-cell T cell populations, which is hindered by label heterogeneity. To resolve this, the authors propose TCellAlignโ€”the first multi-agent workflow that frames population alignment as an evidence-based task. Integrating large language models, ontological resources, and literature mining, TCellAlign standardizes cell type labels while preserving original terminological evidence and adhering to expert naming conventions. Evaluated on a manually curated benchmark encompassing 44 studies and over seven million cells, TCellAlign demonstrates significant improvements over existing ontology-based methods in both semantic consistency and transcriptomic coherence.
๐Ÿ“ Abstract
Cell type standardization plays a central role in integrating biological knowledge across single-cell studies. While standardized resources (e.g., Cell Ontology, Nomenclature Frameworks) provide unified vocabularies of cell populations, scientific publications and public datasets continue to use heterogeneous study-specific labels, making cross-study comparison difficult even when biologically equivalent cell populations are described. In this work, we are the first to formulate this challenge as an evidence-grounded cell population alignment problem and propose TCellAlign, a multi-agent framework that includes literature retrieval, information extraction, nomenclature-guided label alignment, and evidence-based adjudication. This modular design preserves the original terminology and supporting evidence reported by each study while producing standardized labels that can be compared across studies. We further construct a manually validated benchmark dataset linking study-specific labels, CZ CELLxGENE annotations, and standardized T-cell nomenclature across 44 manually curated, published studies (including over seven million cells) spanning four biological categories: healthy, cancer, infectious disease and inflammatory diseases. Across the evaluated tasks, TCellAlign achieves stronger semantic agreement than ontology-based baselines and maintains transcriptomic coherence with both open-source and closed-source large language models (LLM) backbones. By connecting literature, datasets, and expert's nomenclature, TCellAlign enables consistent interpretation of T-cell subtypes and states across studies, facilitating biological knowledge integration and the development of future foundation models built upon standardized cellular representations.
Problem

Research questions and friction points this paper is trying to address.

cell type standardization
cross-study comparison
T-cell nomenclature
heterogeneous labels
single-cell studies
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent workflow
nomenclature-guided alignment
evidence-grounded cell population alignment
T-cell standardization
cross-study integration
P
Pengyu Xie
The Chinese University of Hong Kong, Shenzhen, School of Artificial Intelligence
R
Rongjia Zhou
Emory University, Department of Computer Science and Biology
Z
Zhilin Ou
The Chinese University of Hong Kong, Shenzhen, School of Artificial Intelligence
J
Junyuan Zhang
The University of Melbourne, Department of Biochemistry and Pharmacology
Xiang Zhou
Xiang Zhou
Yale University
StatisticsGeneticsGenomics
X
Xiaobo Sun
Emory University, Department of Human Genetics
Jiaying Lu
Jiaying Lu
Research Assistant Professor of School of Nursing's Center for Data Science, at Emory University
AI for HealthcareKnowledge GraphMultimodal LearningLarge Language Model
W
Wenjing Ma
The Chinese University of Hong Kong, Shenzhen, School of Artificial Intelligence