🤖 AI Summary
Current evaluations of synthetic medical data predominantly rely on statistical similarity and predictive performance, which often fail to capture clinical validity. This work proposes an epidemiology-informed, multidimensional evaluation framework that systematically assesses generative models on structured electronic health records across three dimensions: descriptive fidelity, clinical utility, and structural validity. Applying this framework to a real-world PRIME-CVD cohort comprising 50,000 individuals, the authors empirically compare four model families—GANs, VAE-based approaches, diffusion models, and masked modeling—and find that even models achieving high distributional fidelity exhibit miscalibration and distorted inter-variable relationships. These findings reveal that conventional evaluation metrics tend to overestimate synthetic data quality and underscore the necessity of shifting toward domain-driven validation of clinical effectiveness.
📝 Abstract
Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity. We introduce a multi-dimensional evaluation framework grounded in epidemiology, assessing descriptive fidelity, clinical utility, and structural validity, corresponding to descriptive, predictive, and causal questions. We evaluate four representative generative paradigms - GAN-based, VAE-boosted, diffusion-based, and masked modelling - using PRIME-CVD, a 50,000-person cohort with known ground-truth structure. While all models reproduce marginal distributions, none simultaneously preserve subgroup structure, effect estimates, and dependency structure. Notably, models with strong distributional fidelity can exhibit poor calibration and distorted relationships, leading to unreliable inference. These results show that current evaluation practices can overestimate synthetic data quality and motivate domain-informed assessment based on the ability to support valid clinical and scientific conclusions.