🤖 AI Summary
This study addresses the significant bias of Fréchet Inception Distance (FID) estimators under finite-sample regimes, which constrains the evaluation efficiency of generative models. By analyzing the FID estimation error between Gaussian distributions and establishing tight bias-variance bounds, this work proposes an arbitrary-order extrapolation framework alongside RTD, a debiasing algorithm based on U-statistics. Theoretically, the proposed approach achieves optimal sample complexity. Empirically, RTD attains the lowest estimation error under standard computational budgets, while VALE2 matches the accuracy of conventional 50,000-sample evaluations using only 10,000 samples, thereby substantially reducing computational costs.
📝 Abstract
The Fréchet Inception Distance (FID) is widely used to evaluate generative models, but its empirical plug-in estimator suffers from finite-sample bias [BSAG18, CF20]. We study the sample complexity $n$ of estimating FID to error $ε$ between $d$-dimensional Gaussians with bounded mean distance and covariances, when one distribution is known. Our contributions are threefold. (1) We establish tight finite-sample $Θ(\frac{d^2}{n})$ bias and $Θ(\frac{d}{n} + \frac {d^2} {n^2})$ variance bounds for the empirical plug-in estimator, establishing a $\gtrsim d^2$ sample complexity. (2) To debias the empirical plug-in estimator, we generalize the ${\rm FID}_\infty$ estimator of [CF20] to extrapolation methods of arbitrary order $k$. We further prove tight bias and variance bounds of $Θ(\frac{d^{k + 2}}{n^{k + 1}})$ and $Θ(\frac d n + \frac{d^2}{n^2})$ for any order-$k$ extrapolation under our framework. (3) We introduce Relative Taylor Debiasing (RTD), a new, computationally efficient FID estimation algorithm using debiasing techniques inspired by U-statistics. We show that RTD achieves an $O(\frac d {ε^2})$ sample complexity, and prove that this is optimal. We provide a complementary empirical evaluation of our new estimators. Our experiments on synthetic Gaussians validate the predicted residual bias and support the tightness of our bounds. On ImageNet with Inception embeddings, RTD achieves the lowest mean estimation error at the standard 50K sample budget, while our second-order variance-aware extrapolation estimator (VALE$_2$) uses only 10K samples to achieve accuracy comparable to FID$_\infty$ at 50K samples.