🤖 AI Summary
This work investigates the sensitivity of optimal score functions and generated samples in diffusion models to perturbations in the underlying data distribution—without requiring model retraining. We propose the first differentiable analytical framework that explicitly derives a closed-form expression for the directional derivative of the mapping from data distribution to score function. The method supports black-box access to pretrained models, requiring only their forward outputs and input gradients. By incorporating numerically stable differentiation techniques, it achieves sensitivity estimation with computational complexity matching that of standard sampling. Experiments on image diffusion models demonstrate high-precision prediction of how generated sample distributions respond to minor training-set perturbations. Predicted changes correlate strongly with actual changes observed after retraining or fine-tuning (Pearson *r* > 0.92), validating its fidelity. This framework provides a novel tool for model diagnostics, robustness analysis, and data editing in diffusion-based generative modeling.
📝 Abstract
Training a diffusion model approximates a map from a data distribution $ρ$ to the optimal score function $s_t$ for that distribution. Can we differentiate this map? If we could, then we could predict how the score, and ultimately the model's samples, would change under small perturbations to the training set before committing to costly retraining. We give a closed-form procedure for computing this map's directional derivatives, relying only on black-box access to a pre-trained score model and its derivatives with respect to its inputs. We extend this result to estimate the sensitivity of a diffusion model's samples to additive perturbations of its target measure, with runtime comparable to sampling from a diffusion model and computing log-likelihoods along the sample path. Our method is robust to numerical and approximation error, and the resulting sensitivities correlate with changes in an image diffusion model's samples after retraining and fine-tuning.