scMIR: a vision-language foundation model for single-cell light microscopy image representation

๐Ÿ“… 2026-07-20
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of achieving generalizable analysis of single-cell optical microscopy images, which exhibit high heterogeneity and are sensitive to variations in imaging conditions, thereby hindering cross-cell-type, cross-modality, and cross-experimental comparability. The work proposes a self-supervised foundation model that, for the first time, integrates visionโ€“language co-modeling in this domain. By jointly learning cellular morphology and biological semantics through image reconstruction and text-guided cross-modal alignment, the model constructs a unified representation space. Remarkably, without task-specific fine-tuning, it demonstrates strong generalization across diverse downstream tasks, significantly outperforming both existing general-purpose models and specialized methods on 16 benchmark datasets, encompassing key applications such as cell classification, clustering, phenotype inference, and batch effect correction.
๐Ÿ“ Abstract
Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis. Existing representation learning methods mostly rely on task-oriented modeling, which is limited by specific datasets and predefined tasks, making them difficult to generalize across different cell types and microscopy modalities, and experimental conditions. Although general-purpose methods have improved the generalization ability of image representation in recent years, their limited utilization of experimental background and biological context information still poses challenges in complex phenotypic analysis. Here, we propose scMIR, a vision-language foundation model for single-cell light microscopy image representation. By synergistically combining self-supervised image reconstruction with text-guided cross-modal alignment, scMIR can simultaneously encode morphological and biological semantic information in a unified representation space. scMIR is pre-trained on 207,957 image-text pairs, covering various cell types, microscopy modalities, and perturbation conditions. scMIR outperforms existing general models and task-oriented methods as systematically evaluated on various complex tasks using 16 benchmark datasets, including cell classification, clustering, phenotype inference, and batch effect correction tasks. Furthermore, scMIR shows a strong generalization ability across various tasks without requiring task-specific fine-tuning. With its unique advantages, we envision scMIR may promote the standardization and automation of high-throughput phenotyping workflows through supporting various downstream analysis tasks.
Problem

Research questions and friction points this paper is trying to address.

single-cell microscopy
image representation
phenotypic analysis
generalization
biological context
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language foundation model
single-cell microscopy
cross-modal alignment
self-supervised representation learning
phenotypic generalization
๐Ÿ”Ž Similar Papers
No similar papers found.
Y
Yifan Shang
Laser Metrology and Biomedicine Lab, Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China
J
Jiahui Tan
State Key Laboratory of Chemo and Biosensing, College of Computer Science and Electronic Engineering, Hunan University, Changsha 410082, China
Xiangxiang Zeng
Xiangxiang Zeng
Deparment of Computer Science, Hunan University
Computational IntelligenceAI4ScienceAI for Drug Discovery
Renjie Zhou
Renjie Zhou
The Chinese University of Hong Kong
Quantitative Phase ImagingOptical Diffraction Tomography