🤖 AI Summary
This paper addresses the failure of conventional evaluation metrics for synthetic tabular data under few-shot settings. We demonstrate that global statistical measures—such as Maximum Mean Discrepancy (MMD) and propensity score-based metrics—exhibit unreliable statistical significance and severe misalignment with the underlying distribution’s topological structure when training samples are scarce. To overcome this, we systematically argue for multi-dimensional, synergistic evaluation in few-shot regimes and propose the normalized Bottleneck distance as a novel, topology-aware robustness metric. Experiments across four few-shot benchmarks reveal that mainstream metrics consistently overestimate distributional similarity; while our proposed metric offers complementary structural insights, it exhibits high cross-experiment variability and boundary violations—confirming the unreliability of any single metric. Our core contribution is the formal identification of topological instability as the fundamental challenge in few-shot evaluation, and the establishment of the first integrated verification framework unifying statistical divergence, optimal transport, and topological data analysis (TDA).
📝 Abstract
This work proposes a method to evaluate synthetic tabular data generated to augment small sample datasets. While data augmentation techniques can increase sample counts for machine learning applications, traditional validation approaches fail when applied to extremely limited sample sizes. Our experiments across four datasets reveal significant inconsistencies between global metrics and topological measures, with statistical tests producing unreliable significance values due to insufficient sample sizes. We demonstrate that common metrics like propensity scoring and MMD often suggest similarity where fundamental topological differences exist. Our proposed normalized Bottleneck distance based metric provides complementary insights but suffers from high variability across experimental runs and occasional values exceeding theoretical bounds, showing inherent instability in topological approaches for very small datasets. These findings highlight the critical need for multi-faceted evaluation methodologies when validating synthetic data generated from limited samples, as no single metric reliably captures both distributional and structural similarity.