ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue whereby random partitioning of datasets containing potentially correlated samples yields non-independent training and test sets, leading to overestimated generalization. To mitigate this, we propose a modality-agnostic, relation-aware data partitioning framework. The method constructs a proximity graph based on a hierarchical latent variable model and employs community detection algorithms to identify groups of related samples, thereby generating statistically independent subsets. Furthermore, an unlabeled resolution-adaptive mechanism is introduced to enable automatic adjustment of partition granularity and low-cost estimation of effective dataset size. Experiments on molecular and protein datasets demonstrate that the proposed framework significantly enhances scalability while preserving predictive performance, supporting large-scale relation-aware partitioning and diversity-aware data scaling.
📝 Abstract
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.
Problem

Research questions and friction points this paper is trying to address.

random splitting
latent relations
generalization estimates
dataset scalability
train-test independence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Relations
Scalable Framework
Community Detection
Modality-Agnostic
Dataset Splitting
🔎 Similar Papers
No similar papers found.
A
Anthony Lavertu
Department of Computer Science, Université Laval, Québec, QC, Canada
J
Jacob Cote
Department of Computer Science, Université Laval, Québec, QC, Canada
S
Sophie Gobeil
Department of Biochemistry, Microbiology and Bioinformatics, Université Laval, Québec, QC, Canada
Jacques Corbeil
Jacques Corbeil
Department of Molecular Medicine, Université Laval, Québec, QC, Canada
I
Isabeau Premont-Schwarz
Department of Computer Science, Université Laval, Québec, QC, Canada
Pascal Germain
Pascal Germain
Associate Professor, Université Laval
Machine Learning