Investigation into using stochastic embedding representations for evaluating the trustworthiness of the Fr\'{e}chet Inception Distance

📅 2026-01-29
📈 Citations: 0
Influential: 0
📄 PDF

career value

197K/year
🤖 AI Summary
This work addresses the limited reliability of the Fréchet Inception Distance (FID) in non-natural image domains such as medical imaging, where its dependence on an ImageNet1K-pretrained Inception-v3 model impedes effective representation of out-of-distribution data. To mitigate this issue, the authors propose generating stochastic embedding representations via Monte Carlo Dropout and computing the predictive variance of FID feature embeddings to quantify how far input data deviate from the training distribution. This study introduces predictive variance as a novel indicator of FID reliability, revealing a clear link between embedding uncertainty and distributional shift. Empirical results demonstrate that this variance strongly correlates with the degree of out-of-distributionness on both an augmented ImageNet validation set and external medical imaging datasets, offering a new and robust perspective for evaluating synthetic medical image quality.

Technology Category

Application Category

📝 Abstract
Feature embeddings acquired from pretrained models are widely used in medical applications of deep learning to assess the characteristics of datasets; e.g. to determine the quality of synthetic, generated medical images. The Fr\'{e}chet Inception Distance (FID) is one popular synthetic image quality metric that relies on the assumption that the characteristic features of the data can be detected and encoded by an InceptionV3 model pretrained on ImageNet1K (natural images). While it is widely known that this makes it less effective for applications involving medical images, the extent to which the metric fails to capture meaningful differences in image characteristics is not obviously known. Here, we use Monte Carlo dropout to compute the predictive variance in the FID as well as a supplemental estimate of the predictive variance in the feature embedding model's latent representations. We show that the magnitudes of the predictive variances considered exhibit varying degrees of correlation with the extent to which test inputs (ImageNet1K validation set augmented at various strengths, and other external datasets) are out-of-distribution relative to its training data, providing some insight into the effectiveness of their use as indicators of the trustworthiness of the FID.
Problem

Research questions and friction points this paper is trying to address.

Fréchet Inception Distance
trustworthiness
medical images
out-of-distribution
feature embeddings
Innovation

Methods, ideas, or system contributions that make the work stand out.

stochastic embedding
Fréchet Inception Distance
Monte Carlo dropout
predictive variance
out-of-distribution detection
🔎 Similar Papers
No similar papers found.
C
Ciaran Bench
Department of Data Science and AI, National Physical Laboratory, Teddington, UK
V
Vivek Desai
Department of Data Science and AI, National Physical Laboratory, Teddington, UK
C
Carlijn Roozemond
Dutch Expert Centre for Screening (LRCB), Nijmegen, The Netherlands
R
Ruben E. van Engen
Dutch Expert Centre for Screening (LRCB), Nijmegen, The Netherlands
S
Spencer A. Thomas
Department of Data Science and AI, National Physical Laboratory, Teddington, UK