Score
Designs and implements evaluation pipelines and metric implementations to compute the Fréchet Inception Distance (FID) and its variants (e.g., guided-class FID, child FID) on generated images; builds analyses that detect distribution shift or manifold damage caused by guidance, and benchmarks/compares guidance methods and quality–compute tradeoffs for image generative models.
This work systematically investigates the discrepancy between Fréchet Inception Distance (FID) and human perception, demonstrating that low FID scores do not necessarily correspond to high-quality generated outputs. The study identifies the geometric structure of the reference dataset—particularly its distribution density and effective rank—as the primary cause of FID’s unreliability. Through empirical analyses involving precision-recall decomposition, multiple feature spaces, and ablation studies on distance metrics, the authors reveal that FID exhibits reasonable behavior on concentrated datasets but becomes misleading on dispersed ones. Experiments across six datasets confirm that the geometric properties of the reference data critically influence the fidelity of distribution-based evaluation metrics. The findings advocate for interpreting such metrics in conjunction with the underlying dataset geometry, offering a more reliable foundation for evaluating generative models.
Current evaluations of generative models often report a single FID score, neglecting the randomness inherent in both training and sampling processes, which undermines reliability and reproducibility. This work is the first to model FID as a random variable jointly determined by training and sampling seeds. By training hundreds of SiT models on ImageNet at 256×256 resolution, we systematically quantify the sources of FID variance. Our analysis reveals that training seeds dominate FID variation—contributing approximately 3.2 times more than sampling seeds—and that optimal classifier-free guidance reduces FID dispersion by nearly half. Moreover, FID differences with a coefficient of variation (CoV) below 1.3% should be deemed statistically insignificant. Building on these insights, we propose a new evaluation paradigm incorporating multi-seed error bars and optimal guidance, substantially enhancing robustness and reproducibility in generative model assessment.
This paper identifies a systematic misalignment between generic generative evaluation metrics—such as Fréchet Inception Distance (FID)—and downstream task performance (e.g., classification or segmentation) in retinal image synthesis. Method: Through systematic experiments across multimodal retinal datasets (fundus photography and OCT), we empirically analyze the correlation between FID (and its variants) and actual gains in downstream model performance. Contribution/Results: We provide the first empirical evidence that FID scores fail to predict whether synthetic data meaningfully improve downstream model accuracy. To address this, we propose a “task-driven evaluation” paradigm, advocating direct assessment via target downstream task performance—replacing proxy metrics reliant on ImageNet-pretrained features. Our findings are robustly replicated across multiple public retinal image benchmarks, offering both methodological insight and practical guidance for evaluating biomedical image generation.
提出基于小波包变换的Fréchet Wavelet Distance(FWD)指标,解决现有生成图像评价指标对特定生成器和数据集的偏见问题,通过计算小波包系数空间的Fréchet距离,实现领域无关且更可解释的质量评估。
This work addresses a critical limitation of Training-Free Guidance (TFG) in high-noise regimes, where estimating a clean image from pure noise often causes the guidance signal to deviate from the data manifold, degrading generation quality. The study is the first to reveal that the choice of prediction target—ε, v, or x—profoundly influences TFG’s ability to preserve the data manifold. Theoretical analysis demonstrates that x-prediction directly yields the clean image, substantially reducing estimation error under high noise. To detect such manifold distortions overlooked by conventional metrics, the authors introduce a novel evaluation measure, guided-class FID (Child FID). Experiments on a newly curated fine-grained bird benchmark and style transfer tasks confirm that TFG with x-prediction significantly outperforms other prediction strategies in maintaining sample manifold consistency.
This work addresses critical limitations of existing generative model evaluation metrics—such as the Fréchet Inception Distance (FID)—in sample complexity, computational efficiency, and robustness against adversarial attacks. The authors propose the Monge Inception Distance (MIND), which computes a one-dimensional optimal transport average of Inception features via sliced Wasserstein distance, thereby circumventing the need for high-dimensional mean and covariance estimation. MIND achieves evaluation performance comparable to FID with only 1k–5k samples, offering an order-of-magnitude improvement in sample efficiency and two orders of magnitude faster computation. Moreover, it demonstrates significantly enhanced robustness against adversarial manipulations such as moment matching. While maintaining high correlation with FID, MIND exhibits superior discriminative power, substantially facilitating efficient generative model development and iteration.
This work addresses the issue of “Fréchet hacking” in generative model optimization, where Fréchet distance losses based on static pretrained feature spaces yield deceptively high scores despite degraded visual quality and poor cross-feature alignment. To mitigate this, the authors propose the adversarial Fréchet distance (AdvFD) loss, which introduces adversarial learning into Fréchet distance optimization for the first time. AdvFD constructs a learnable, adaptive feature space that dynamically enhances distribution discrepancy measurement and incorporates a real-feature whitening mechanism to suppress feature amplification and stabilize training. The method consistently improves both visual fidelity and distribution alignment in single-step generator post-training across various model scales and backbone architectures, including JiT and pMF.
Existing generative models often suffer from mode collapse when optimizing the Fréchet Inception Distance (FID), struggling to balance sample quality and diversity. This work proposes a Recursive Tuning Mechanism (RTM), which, for the first time, integrates iterative latent refinement into the generative process, replacing the single-pass forward mapping used in architectures like StyleGAN2. RTM significantly enhances mode coverage without compromising fidelity. By combining implicit maximum likelihood estimation (IMLE) with a precision-recall evaluation framework, RTM achieves state-of-the-art precision and recall on CIFAR-10, CelebA-HQ, and nine few-shot benchmarks while maintaining excellent FID scores. Moreover, it substantially outperforms StyleGAN2 and its ADA variant on high-resolution tasks such as AFHQ-v1.
This work addresses a critical issue in high-dimensional flow matching: the systematic underestimation of velocity magnitude at trajectory initialization induces integration lag, preventing generated samples from accurately reaching the data manifold. The study uncovers an asymmetric mechanism wherein velocity contraction is detrimental at the start but beneficial near the end of trajectories. To mitigate this without retraining, the authors propose a joint strategy combining a Scale Scheduling Corrector (SSC) with Magnitude-Aware Flow Matching (MAFM), implementable with just a single line of code. The approach substantially improves performance—on ImageNet-1k, it reduces FID from 13.68 to 7.58 (a 44.6% improvement), achieves a 5× speedup, and yields 50-step generation quality surpassing the original 250-step baseline; on MS-COCO text-to-image generation, FID also improves by approximately 22%.