Language as the Interface: Foundation-Model Contrastive Learning Links Transcriptomes and Electrophysiology

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of multimodal integration between transcriptomic and electrophysiological data in neuroscience, particularly the difficulties of cross-modal alignment and cross-species transfer. To this end, it proposes LangPatch, a framework that introduces a novel cross-modal contrastive learning paradigm built upon a frozen text encoder. By integrating GenePT pretrained representations with context adapters and projection modules, the method achieves efficient alignment between gene expression profiles and neuronal electrophysiological features through a language interface. This work establishes a multimodal foundation model that unifies molecular and functional representations. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines in predictive performance and successfully enables cross-domain generalization from mouse models to the human brain.
📝 Abstract
Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides paired measurements of gene expression and intrinsic electrophysiology from the same neuron, establishing a basis for training cross-modal models. Here we introduce LangPatch, a foundation-model-based contrastive learning framework that uses paired Patch-seq data to align pretrained GenePT representations with electrophysiological phenotypes through a language-based interface. Gene descriptions and verbalized electrophysiological profiles are embedded by the same frozen text encoder. A context adapter and projection modules connect the modalities through paired contrastive learning. Across mouse visual, mouse motor, and human cortical cohorts, LangPatch achieves the highest mean transcriptome-to-electrophysiology prediction correlation among the evaluated foundation-model and representation-learning methods. It also improves held-out cross-modal alignment in the two mouse cohorts (FOSCTTM 0.107/0.135 vs. 0.208/0.222 for JAMIE, an existing cross-modal Patch-seq imputation method). It predicts transcriptomic family, type, cortical layer, and marker-gene expression from electrophysiology, exceeding other baselines on most endpoints. More importantly, the method transfers across brain areas and species: a model trained on mouse visual cortex predicts electrophysiology in motor cortex with approximately 70% correlation retention and in human cortex with 47% (58% on acute-slice recordings). Together, these results demonstrate alignment between molecular and functional representations of neurons, providing a building block for multimodal foundation models in neuroscience.
Problem

Research questions and friction points this paper is trying to address.

transcriptomics
electrophysiology
cross-modal alignment
multimodal foundation models
Patch-seq
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Learning
Foundation Model
Cross-modal Alignment
Patch-seq
Language Interface
💼 Related Jobs
No related jobs found.
J
Junbo Shen
Department of Computer Science and Engineering, The Chinese University of Hong Kong
J
Jinying Gao
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Bo Lei
Bo Lei
Beijing Academy of Artificial Intelligence
NeuroscienceArtificial intelligence