BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Single-cell transcriptomic data are highly sparse, and conventional self-supervised methods that reconstruct individual genes struggle to capture coordinated gene regulation. This work proposes a Joint Embedding Predictive Architecture (JEPA) based on graph-connected gene blocks, which, for the first time, treats functional gene modules—defined by protein–protein interactions and co-expression—as prediction units. In this framework, a student network infers representations of target gene blocks from partially observed inputs, while a teacher network provides target representations derived from the full gene set, integrated with linear attention and knowledge distillation strategies. The approach substantially improves embedding effective rank, reduces dependence on sequencing depth, and preserves biological pathway and perturbation response information. It achieves the lowest perturbation response error on CellBench and, compared to scFoundation, yields 5.75× and 3.76× improvements in fine-tuning accuracy and embedding throughput, respectively, on the hPancreas task.
📝 Abstract
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.
Problem

Research questions and friction points this paper is trying to address.

single-cell transcriptomics
self-supervised learning
gene co-regulation
representation learning
data sparsity
Innovation

Methods, ideas, or system contributions that make the work stand out.

joint-embedding prediction
graph-connected gene blocks
self-supervised learning
linear attention
single-cell transcriptomics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs