ARO: Aligned Representation learning for multi-Omics data

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses critical bottlenecks in cancer multi-omics analysis, including missing data, unmatched samples, and high assay costs, by proposing the ARO model. Leveraging representation learning and multimodal alignment techniques, ARO constructs a unified representation space for multi-omics data. Unlike large foundation models that require massive datasets, ARO prioritizes clinical practicality, enabling high-fidelity reconstruction of missing modalities and robust downstream classification. Experimental results demonstrate that the model achieves a mean squared error (MSE) of 0.15 on the validation set, effectively facilitating multi-omics integrative inference under data-constrained scenarios. By significantly reducing reliance on large-scale experimental profiling, this work presents an efficient and practical paradigm for incomplete multi-omics analysis.
📝 Abstract
The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities. Contrary to increasingly complex, larger models, e.g. Foundation Models (FMs), ARO prioritizes practical applicability in limited or incomplete data settings. ARO optimally reconstructs missing modalities (MSE of $0.15$ on the validation and test data in the Unmasked settings), with its learned latent embeddings enabling a downstream cancer classification task. Our findings indicate that analyzing diverse molecular layers as a single integrated system offers a reliable and cost-efficient approach, reducing dependence on large-scale experimental testing, while still supporting multi-omic exploration in limited data settings.
Problem

Research questions and friction points this paper is trying to address.

multi-omics data
missing modalities
representation learning
computational biology
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-omics
Representation Learning
Missing Modality Reconstruction
Aligned Embeddings
Cancer Classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Amogh Singh
Department of Computer Science, University of Cambridge, Cambridge, UK
Y
Yash Shah
Department of Computer Science, University of Cambridge, Cambridge, UK
C
Chiara D'Ercoli
DIAG, University of Rome “Sapienza”, Rome, Italy
Arash Mehrjou
Arash Mehrjou
ETH Zürich - Max Planck Institute - GSK.ai
Machine LearningControl TheoryCausality
Patrick Schwab
Patrick Schwab
GSK
Causal Machine LearningAI in Drug DiscoveryAI in HealthcareAI in Medicine
T
Timothy Jones
Department of Computer Science, University of Cambridge, Cambridge, UK
Pietro Liò
Pietro Liò
Professor, University of Cambridge
AI & Comp Biology -> Medicine