🤖 AI Summary
This study addresses the issue whereby random partitioning of datasets containing potentially correlated samples yields non-independent training and test sets, leading to overestimated generalization. To mitigate this, we propose a modality-agnostic, relation-aware data partitioning framework. The method constructs a proximity graph based on a hierarchical latent variable model and employs community detection algorithms to identify groups of related samples, thereby generating statistically independent subsets. Furthermore, an unlabeled resolution-adaptive mechanism is introduced to enable automatic adjustment of partition granularity and low-cost estimation of effective dataset size. Experiments on molecular and protein datasets demonstrate that the proposed framework significantly enhances scalability while preserving predictive performance, supporting large-scale relation-aware partitioning and diversity-aware data scaling.
📝 Abstract
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.