🤖 AI Summary
To address the challenge of integrating multi-omics data under few-shot, unpaired, and weakly supervised settings, this paper proposes the Semi-Supervised Probabilistic Coupling Framework (SPCF), enabling cross-sample alignment of heterogeneous modalities within a shared latent space. SPCF integrates probabilistic graphical modeling with variational inference, introducing the first semi-supervised alignment mechanism capable of operating with extremely limited supervision—fewer than ten paired samples—or even in fully unpaired scenarios. Through controlled synthetic experiments, we quantitatively characterize the minimal supervision required, facilitating rapid generalization to novel disease contexts. Under sparse labeling, SPCF significantly improves modality alignment accuracy and boosts downstream classification and anomaly detection performance by an average of 12.6% (p < 0.01). The implementation is publicly available.
📝 Abstract
A key challenge today lies in the ability to efficiently handle multi-omics data since such multimodal data may provide a more comprehensive overview of the underlying processes in a system. Yet it comes with challenges: multi-omics data are most often unpaired and only partially labeled, moreover only small amounts of data are available in some situation such as rare diseases. We propose MODIS which stands for Multi-Omics Data Integration for Small and unpaired datasets, a semi supervised approach to account for these particular settings. MODIS learns a probabilistic coupling of heterogeneous data modalities and learns a shared latent space where modalities are aligned. We rely on artificial data to build controlled experiments to explore how much supervision is needed for an accurate alignment of modalities, and how our approach enables dealing with new conditions for which few data are available. The code is available athttps://github.com/VILLOUTREIXLab/MODIS.