🤖 AI Summary
This study addresses the discrepancy between the empirical performance of private evolutionary algorithms, which often exceeds theoretical predictions, and the convergence difficulties encountered by standard variants. To resolve this, the work reformulates private evolution as a generative model-augmented Wasserstein learning framework that leverages intrinsic dimensionality to optimize sample complexity. It elucidates the theoretical gains of generative models in distribution capture and introduces a geometry-aware variant to ensure convergence. By integrating differential privacy, Wasserstein learning, and generative modeling, the proposed algorithm significantly improves recall across benchmark evaluations while achieving stable convergence. Ultimately, this research delivers a method that combines rigorous theoretical guarantees with strong empirical competitiveness.
📝 Abstract
Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.