Score
Designs and implements noise schedules that vary the magnitude or timing of noise applied per class according to each class's empirical frequency, so training and sampling treat rare classes differently from common ones. Builds class-aware noise-magnitude or scheduling rules and integrates them into score-based / diffusion model training and sampling to reduce score-estimation errors in low-density regions and improve diversity and quality of generated samples for infrequent classes.
This work investigates the design principles of noise schedules in diffusion models and their impact on sampling quality and training stability. We propose the first unified classification framework for noise schedules—encompassing linear, cosine, and learned variants—and systematically analyze their properties via theoretical derivation, controlled ablation experiments, and visualization-based evaluation. Our analysis reveals a fundamental trade-off: optimal noise scheduling must jointly optimize signal-to-noise ratio (SNR) smoothness and denoising gradient stability. Based on this insight, we derive practical, interpretable design principles that balance computational efficiency and generation fidelity. The resulting guidelines provide theoretically grounded, reproducible support for developing high-performance diffusion samplers.
This work addresses the degradation in generation quality and diversity for low-frequency classes in diffusion models, which arises from biased score estimation due to data sparsity, while high-frequency classes dominate the score space and exacerbate class imbalance under long-tailed distributions. The study is the first to uncover the relationship between class frequency and multi-scale noise scheduling, proposing a class-frequency-guided noise scheduling strategy that assigns larger noise scales to low-frequency classes to improve their score estimation. Evaluated on long-tailed benchmarks such as CIFAR-100-LT and ImageNet-LT, the method significantly enhances generation quality and diversity across all classes, outperforming existing baselines in image generation, classification, and text-to-image tasks, thereby demonstrating the effectiveness and generalizability of frequency-aware noise scheduling.
This work addresses the lack of theoretical grounding in noise scheduling for diffusion models, which has hindered the understanding of empirically effective scheduling strategies. The authors formulate the problem for the first time as an optimal control problem, where the Fisher information serves as the state variable and the noise schedule acts as the control input, with the objective of minimizing an upper bound on the KL sampling error. Within this framework, they derive sufficient conditions for achieving near-optimal sampling error and obtain a tunable closed-form expression for the noise schedule. The proposed method unifies and generalizes exponential and sigmoidal schedules, and, after parameter tuning, achieves improved Fréchet Inception Distance (FID) scores on standard image generation benchmarks.
Traditional diffusion models rely on handcrafted noise schedules, which often waste computational resources in low-information regions and exhibit poor generalization. This work proposes InfoNoise, the first noise scheduling mechanism grounded in information theory that adapts to data by dynamically guiding the noise sampling distribution through estimation of the conditional entropy rate in the forward process, thereby replacing heuristic design. Requiring no manual hyperparameter tuning, InfoNoise matches or surpasses the performance of carefully tuned EDM schedules on natural images and accelerates CIFAR-10 training by 1.4×. On discrete data, it achieves superior generation quality with up to three times fewer training steps.
This work addresses the limitations of conventional diffusion models, which rely on handcrafted noise schedules that struggle to balance efficiency and generation quality across varying resolutions and often involve redundant steps. The authors propose a spectrum-guided, instance-adaptive noise scheduling method that derives a “tight” noise schedule grounded in the spectral characteristics of images and theoretical bounds of the diffusion process. During inference, this schedule is dynamically applied through conditional sampling. By incorporating spectral information into the selection of noise levels—a first in the field—the approach significantly enhances generation quality at low inference step counts within single-stage, pixel-level diffusion models while reducing unnecessary computation.
Existing diffusion models rely heavily on heuristic, empirically designed noise schedules lacking theoretical grounding, leading to spectral discrepancies between generated and real data distributions. This work introduces the first spectral-analysis-based framework for noise schedule design: it formalizes diffusion sampling as a closed-form frequency-domain transfer function, enabling theoretically principled alignment between the noise schedule and the intrinsic spectral characteristics of data. Under assumptions of Gaussianity and translation invariance, we derive an analytical frequency-domain response model and develop a data-driven, adaptive schedule optimization algorithm. The resulting frequency-aware noise schedules significantly improve sampling efficiency—accelerating generation by up to 1.8×—and enhance sample quality, reducing FID by 12.3%. Our approach establishes a mathematically interpretable foundation for noise scheduling and provides a practical, spectrum-informed design paradigm for diffusion models.
This study addresses the sensitivity to noise schedules and the lack of theoretical justification in multi-step sampling for consistency models. By analyzing the composition of noising and denoising operators, it establishes a non-asymptotic convergence theory under explicitly verifiable stability assumptions. Methodologically, the analysis decouples initialization error contraction from approximation error accumulation, revealing that large early-stage noise drives contraction while small late-stage noise controls residual bias, with explicit constants derived for strongly log-concave targets. Experiments confirm that the theoretically predicted contraction and approximation profiles are reliably measurable. This work provides both rigorous theoretical guidance and a practical framework for designing multi-step consistency samplers.
该研究通过全局调度去噪轨迹,优化扩散模型在推理时的计算资源分配,提高样本质量,减少函数评估次数。
This work addresses the limitation of existing hybrid training methods that rely on fixed mixing strategies, which struggle to adapt to the dynamically varying noise present in supervised fine-tuning and reinforcement learning signals. To overcome this, the paper proposes Gradient Adaptive Control (GAC), a method that dynamically adjusts the mixing weights by online estimation of gradient variance and inconsistency between the two signal types. GAC incorporates a noise-aware controller augmented with smoothing mechanisms, prior guidance, and bounded updates, achieving adaptive optimization with negligible computational overhead (<1%). Experimental results demonstrate that GAC consistently outperforms strong baselines across benchmarks in mathematics, code generation, scientific reasoning, and logical tasks, with performance gains becoming more pronounced as model scale increases.
This work addresses the inefficiency of traditional classifier-guided diffusion models, which require separate training of a classifier and a generative model, leading to redundancy and high computational costs. To overcome this, the authors propose an efficient approach for conditional speech generation that repurposes a pretrained speech classifier as a shared backbone network. By freezing the backbone’s parameters and training only lightweight auxiliary subnetworks, the method enables conditional generation while unifying discriminative modeling and speech synthesis within a single architecture for the first time. The framework integrates noise-conditional classification, log-Mel spectrogram space modeling, and denoising score matching, achieving high-quality speech synthesis with substantially reduced memory and computational overhead.
研究通过引入随机化遍历重放方法,解决了类平衡重放中的排练间隔问题,从而在不冻结类共现的情况下控制了排练间隔。