Score
Designs and evaluates methods that probe denoisers and diffusion models by applying controlled frequency‑domain perturbations and measuring how the denoising process redistributes energy across spectral bands; from those measurements derives stable per‑model spectral signatures. Builds analysis and classification procedures that use those spectral fingerprints to attribute individual generated or denoised samples to their source denoiser or diffusion model.
Existing methods struggle to reliably trace generated images back to their source diffusion models, particularly when model architectures or autoencoders are similar. This work proposes a non-intrusive fingerprinting approach that leverages, for the first time, the energy redistribution characteristics of the denoising score function in the spatial frequency domain to construct an intrinsic model signature—requiring no registration, inversion, or optimization. By applying frequency-controllable perturbation probes, the method extracts spectral geometric features from denoising responses using only standard forward passes, yielding Spectral Denoising Signatures (SDS). Evaluated across eight diffusion models, the approach achieves 99.9% source attribution accuracy and maintains robust performance under cross-domain prompt transfer, attaining 96.2% accuracy—significantly outperforming current non-intrusive alternatives.
Existing diffusion models fully degrade data into white noise, resulting in challenging denoising tasks, high computational overhead, and severe loss of spectral structure. To address this, we propose InSPECT—a diffusion model that explicitly preserves invariant spectral features throughout both forward and reverse processes. InSPECT is the first to systematically model and constrain Fourier-domain dynamics, enabling spectral coefficients to smoothly converge toward controllable stochastic noise while jointly preserving structural fidelity, generation diversity, and randomness. Our approach comprises a spectral-aware forward process and a dedicated Fourier-domain denoising network. Evaluated on CIFAR-10, CelebA, and LSUN, InSPECT achieves a 39.23% reduction in FID and a 45.80% improvement in Inception Score (IS) over DDPM, with accelerated convergence and significantly smoother diffusion trajectories.
Existing diffusion models rely heavily on heuristic, empirically designed noise schedules lacking theoretical grounding, leading to spectral discrepancies between generated and real data distributions. This work introduces the first spectral-analysis-based framework for noise schedule design: it formalizes diffusion sampling as a closed-form frequency-domain transfer function, enabling theoretically principled alignment between the noise schedule and the intrinsic spectral characteristics of data. Under assumptions of Gaussianity and translation invariance, we derive an analytical frequency-domain response model and develop a data-driven, adaptive schedule optimization algorithm. The resulting frequency-aware noise schedules significantly improve sampling efficiency—accelerating generation by up to 1.8×—and enhance sample quality, reducing FID by 12.3%. Our approach establishes a mathematically interpretable foundation for noise scheduling and provides a practical, spectrum-informed design paradigm for diffusion models.
To address the degradation in generation quality caused by misalignment between noise levels and distances to the data manifold during diffusion model denoising, this paper proposes a noise-level calibration mechanism. It explicitly models noise level as a proxy for the distance from a sample to the data manifold and introduces a lightweight, plug-and-play auxiliary correction network that dynamically refines denoising estimates at each step. The method requires no retraining of the backbone diffusion model and uniformly supports diverse restoration tasks—including image inpainting, super-resolution, deblurring, color enhancement, and compression artifact removal—while remaining compatible with mainstream samplers (e.g., DDIM). Correction subnetworks are constructed solely from pre-trained denoising networks, incorporating manifold geometric priors and task-specific constraints (e.g., masks, degradation kernels, frequency-domain restrictions). Experiments demonstrate consistent improvements in FID and LPIPS across both unconditional generation and restoration tasks, with minimal computational overhead and stable performance gains.
Existing diffusion models typically employ pointwise reconstruction losses that are insensitive to signal spectra and multiscale structures, often yielding samples with imbalanced frequency content and insufficient structural detail. To address this limitation, this work proposes a lightweight, plug-and-play spectral regularization framework that introduces differentiable Fourier- and wavelet-domain losses as soft inductive biases during standard training. The approach requires no modifications to the diffusion process, model architecture, or sampling procedure, and is compatible with mainstream paradigms such as DDPM, DDIM, and EDM. Evaluated on high-resolution unconditional generation tasks for both images and audio, the method consistently enhances sample quality—particularly in terms of spectral fidelity and multiscale coherence—while incurring negligible computational overhead.
Diffusion models trained on noisy data often produce images with high-frequency artifacts that degrade visual quality. To address this, this work proposes SCoRe, a training-free inference-time method that combines spectral truncation with SDEdit to suppress noise-induced artifacts and regenerate high-frequency details. The key innovation lies in establishing, for the first time, a theoretical mapping between the spectral truncation frequency and the SDEdit initialization timestep, enabling effective noise suppression without retraining. Guided by radial average power spectral density (RAPSD) analysis, SCoRe achieves high-quality image restoration through controlled attenuation and regeneration of high-frequency components. Experiments demonstrate that SCoRe significantly outperforms existing post-processing and noise-robust baselines on CIFAR-10 and SIDD, yielding generations that closely align with the distribution of clean data.
This work addresses the discrepancy in diffusion models between the denoising score matching objective and the Fokker-Planck (FP) equation that governs the true data evolution. While existing strong FP regularization methods are computationally expensive and do not consistently improve generation quality, this study systematically evaluates lightweight regularization terms targeting FP residuals. The authors propose a low-overhead weak FP regularization strategy that substantially reduces computational burden while preserving high-fidelity image generation. Empirical results demonstrate that such lightweight regularization effectively balances efficiency and performance, confirming its feasibility and advantages in practical diffusion model training.
This study addresses the limitation of existing evaluation metrics in revealing the underlying mechanistic differences among generative denoising models. To this end, it proposes the Jacobian spectrum as a key analytical tool for distinguishing such models and establishes a novel training paradigm that directly modulates spectral properties. Specifically, by introducing a spectral regularization method based on perturbed inputs, the approach optimizes responses along data-dependent principal directions while suppressing noise-oriented components. Experiments conducted on ImageNet with both diffusion and flow matching models demonstrate that the proposed method effectively enhances the representation of data-dependent features, leading to substantial improvements in image generation quality.
This work investigates the fundamental reason why Denoising Diffusion Implicit Models (DDIM) are more prone to hallucination—i.e., generating out-of-distribution samples—compared to Denoising Diffusion Probabilistic Models (DDPM). By theoretically analyzing the deterministic ordinary differential equation (ODE) of DDIM and the stochastic differential equation (SDE) reverse dynamics of DDPM under a Gaussian mixture target distribution, the study establishes that DDIM trajectories tend to become trapped along line segments between modes after a critical time, whereas the inherent stochasticity in DDPM enables escape from such regions, thereby suppressing hallucinations. Building on this insight, the authors propose a novel strategy that introduces stochastic steps into the DDIM sampling process. Numerical experiments confirm that DDPM exhibits significantly lower hallucination rates in mode-connecting regions and demonstrate that the proposed method effectively mitigates hallucination in DDIM.
This study addresses the limitation of existing spectral diffusion models, where noise schedules neglect spectral energy distributions, thereby constraining generation efficiency. We propose an energy-adaptive noise scheduling and whitening strategy that, for the first time, introduces image-dependent energy trajectories into the noise schedule. By optimizing the forward diffusion process through global spectral whitening and energy-conditioned noise allocation, our method achieves spectral awareness while preserving Gaussian transition properties. Built upon a transform-domain diffusion framework, the proposed approach is fully compatible with standard DDPM/DDIM architectures. Experiments on CIFAR-10 demonstrate that the Fréchet Inception Distance (FID) improves significantly from 142.48 to 100.45, confirming that our strategy effectively enhances the generative quality of spectral diffusion models.