🤖 AI Summary
This work addresses the degradation of model generalization in supervised learning caused by heterogeneity in training data. To mitigate this issue, the authors propose an input-space adaptive partitioning method grounded in the intrinsic heterogeneity of the data. By introducing a variance-based metric that quantifies the inconsistency in pairwise sample influence, they demonstrate that this variance is maximized under mixture distributions. Leveraging this property, the method automatically partitions the data into homogeneous subsets without requiring prior knowledge, enabling independent training of submodels on each subset. Experiments on EMNIST and synthetic datasets show significant improvements in test accuracy, confirming that the proposed variance metric effectively captures data heterogeneity and offers a novel pathway to enhance model generalization.
📝 Abstract
In this article the authors develop an intrinsic measure for quantifying heterogeneity in training data for supervised learning. This measure is the variance of a random variable which factors through the influences of pairs of training points. The variance is shown to capture data heterogeneity and can thus be used to assess if a sample is a mixture of distributions. The authors prove that the data itself contains key information that supports a partitioning into blocks. Several proof of concept studies are provided that quantify the connection between variance and heterogeneity for EMNIST image data and synthetic data. The authors establish that variance is maximal for equal mixes of distributions, and detail how variance-based data purification followed by conventional training over blocks can lead to significant increases in test accuracy.