compute frechet inception distance

Implement the Fréchet Inception Distance (FID) computation by extracting feature activations from a pretrained Inception network for two image sets, estimating each set’s multivariate Gaussian mean and covariance, and computing the Fréchet (Wasserstein-2) distance between those Gaussians to quantify sample fidelity and diversity; include robust numerical procedures for covariance estimation and matrix square-root operations and report interpretable FID scores.

computefrechetinceptiondistance

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work systematically investigates the discrepancy between Fréchet Inception Distance (FID) and human perception, demonstrating that low FID scores do not necessarily correspond to high-quality generated outputs. The study identifies the geometric structure of the reference dataset—particularly its distribution density and effective rank—as the primary cause of FID’s unreliability. Through empirical analyses involving precision-recall decomposition, multiple feature spaces, and ablation studies on distance metrics, the authors reveal that FID exhibits reasonable behavior on concentrated datasets but becomes misleading on dispersed ones. Experiments across six datasets confirm that the geometric properties of the reference data critically influence the fidelity of distribution-based evaluation metrics. The findings advocate for interpreting such metrics in conjunction with the underlying dataset geometry, offering a more reliable foundation for evaluating generative models.

distributional metricsFréchet Inception Distanceimage generation evaluation

This paper identifies a systematic misalignment between generic generative evaluation metrics—such as Fréchet Inception Distance (FID)—and downstream task performance (e.g., classification or segmentation) in retinal image synthesis. Method: Through systematic experiments across multimodal retinal datasets (fundus photography and OCT), we empirically analyze the correlation between FID (and its variants) and actual gains in downstream model performance. Contribution/Results: We provide the first empirical evidence that FID scores fail to predict whether synthetic data meaningfully improve downstream model accuracy. To address this, we propose a “task-driven evaluation” paradigm, advocating direct assessment via target downstream task performance—replacing proxy metrics reliant on ImageNet-pretrained features. Our findings are robustly replicated across multiple public retinal image benchmarks, offering both methodological insight and practical guidance for evaluating biomedical image generation.

Assessing generative models via downstream tasks in biomedicineEvaluating retinal image synthesis models using FID limitationsMisalignment between FID metrics and biomedical task performance

Current evaluations of generative models often report a single FID score, neglecting the randomness inherent in both training and sampling processes, which undermines reliability and reproducibility. This work is the first to model FID as a random variable jointly determined by training and sampling seeds. By training hundreds of SiT models on ImageNet at 256×256 resolution, we systematically quantify the sources of FID variance. Our analysis reveals that training seeds dominate FID variation—contributing approximately 3.2 times more than sampling seeds—and that optimal classifier-free guidance reduces FID dispersion by nearly half. Moreover, FID differences with a coefficient of variation (CoV) below 1.3% should be deemed statistically insignificant. Building on these insights, we propose a new evaluation paradigm incorporating multi-seed error bars and optimal guidance, substantially enhancing robustness and reproducibility in generative model assessment.

evaluation randomnessFIDgenerative models

This work addresses the limited reliability of the Fréchet Inception Distance (FID) in non-natural image domains such as medical imaging, where its dependence on an ImageNet1K-pretrained Inception-v3 model impedes effective representation of out-of-distribution data. To mitigate this issue, the authors propose generating stochastic embedding representations via Monte Carlo Dropout and computing the predictive variance of FID feature embeddings to quantify how far input data deviate from the training distribution. This study introduces predictive variance as a novel indicator of FID reliability, revealing a clear link between embedding uncertainty and distributional shift. Empirical results demonstrate that this variance strongly correlates with the degree of out-of-distributionness on both an augmented ImageNet validation set and external medical imaging datasets, offering a new and robust perspective for evaluating synthetic medical image quality.

feature embeddingsFréchet Inception Distancemedical images

Fr'echet Wavelet Distance: A Domain-Agnostic Metric for Image Generation

Dec 23, 2023
LV
Lokesh Veeramacheneni
🏛️ University of Bonn

提出基于小波包变换的Fréchet Wavelet Distance(FWD)指标,解决现有生成图像评价指标对特定生成器和数据集的偏见问题,通过计算小波包系数空间的Fréchet距离,实现领域无关且更可解释的质量评估。

Addresses biases in existing metrics like FID and FD-DINOv2Enhances interpretability and robustness across diverse datasetsProposes a domain-agnostic metric for image generation evaluation

Latest Papers

What's happening recently
View more

This work addresses the issue of “Fréchet hacking” in generative model optimization, where Fréchet distance losses based on static pretrained feature spaces yield deceptively high scores despite degraded visual quality and poor cross-feature alignment. To mitigate this, the authors propose the adversarial Fréchet distance (AdvFD) loss, which introduces adversarial learning into Fréchet distance optimization for the first time. AdvFD constructs a learnable, adaptive feature space that dynamically enhances distribution discrepancy measurement and incorporates a real-feature whitening mechanism to suppress feature amplification and stabilize training. The method consistently improves both visual fidelity and distribution alignment in single-step generator post-training across various model scales and backbone architectures, including JiT and pMF.

Fréchet distanceFréchet hackinggenerator post-training

This work addresses the mismatch between the token-level cross-entropy training objective and image-level distribution quality evaluation in autoregressive image generation. The authors propose FD-loss, a post-training method that, for the first time, directly optimizes discrete autoregressive models using the image-level Fréchet distance. By introducing a dual-channel mechanism to construct gradient-free rollout contexts and employing a probability-level straight-through estimator for differentiable replay, the approach updates only the generator without adding parameters or inference overhead. Evaluated on ImageNet at 256×256 resolution, the method reduces the average FID and FD₆ by 41.4% and 52.0%, respectively, and improves the best FID from 2.42 to 1.43, effectively bridging the gap between training objectives and evaluation metrics.

autoregressive image generationcontext mismatchdistributional quality

This work addresses the high computational cost and inefficiency of traditional pixel-based image representations in modeling shape and texture, which stem from high-dimensional discretization. To overcome these limitations, the authors propose a functional data representation that treats image contours and textures as observations of continuous random functions defined over star-shaped domains. By employing star-shaped parameterization, the method unifies the characterization of both shape and texture while avoiding high-dimensional discretization. Integrating functional data analysis with supervised learning, the proposed framework substantially reduces data dimensionality and enhances computational efficiency. Empirical evaluation on real-world image classification tasks demonstrates its effectiveness, confirming that the approach achieves competitive performance with significantly lower computational overhead.

computational costfunctional datahigh-dimensional representation

Hot Scholars

FL

Frantzeska Lavda

University of Geneva
Machine LearningBayesian InferenceNeural Networks
ML

Massimiliano Luca

Researcher @ Fondazione Bruno Kessler
Human MobilityHuman BehaviorComputational Social ScienceCity Science
AK

Alexandros Kalousis

University of Applied Sciences, Western Switzerland (HES-SO).
Machine LearningData Mining
MB

Mukaffi Bin Moin

Department of CSE, Ahsanullah University of Science and Technology
NLP for Social GoodComputer visionLarge Language ModelsTrustworthy AI
KL

Karim Lekadir

ICREA Research Professor, Universitat de Barcelona
Biomedical data sciencehealthcare AItrustworthy AImedical image analysis