🤖 AI Summary
This work addresses the limitation of the traditional Fréchet distance, which evaluates generative models via a single scalar that obscures the specific causes of distributional discrepancies and disconnects the metric from perceptual quality. We introduce the Directional Fréchet Distance, which decomposes optimal transport displacements through projection and leverages multimodal embeddings such as CLIP to map abstract distances onto interpretable semantic directions, revealing that a few key dimensions dominate most deviations. This approach successfully resolves conflicts between FID and human preferences in diffusion models, quantifies frame-wise appearance discrepancies in video FVD, and reinterprets the physical meaning of protein FID. By enabling fine-grained attribution analysis of generative quality across domains, this method provides a principled diagnostic tool for generative modeling. The code is publicly available.
📝 Abstract
The Fr\'echet distance is a de facto standard for evaluating generative models across domains, appearing as FID for images and FVD for videos. It summarizes the discrepancy between generated and reference distributions in a single scalar, with lower values typically interpreted as better generation quality. However, this scalar view can obscure what drives the comparison. For example, in COCO dataset, increasing the number of diffusion sampling steps improves ImageReward scores yet worsens (increases) FID. Motivated by this mismatch, we seek to make the Fr\'echet distance more interpretable by uncovering where the discrepancy lies. To this end, we introduce directional Fr\'echet distance, the expected squared projection of the optimal transport displacement onto a given direction. Across our image, video, and protein case studies, we find that a small number of interpretable directions account for much of the distance. We use these directions to explain the FID increase in terms of semantic concepts represented by CLIP embeddings, quantify FVD's bias toward per-frame appearance, and revisit the interpretation of Protein FID. We open-source our codebase at https://github.com/yhlee-add/directional-fd.