Score
Designs and implements optimization procedures that iteratively modify latent representations of pre-trained generative models by using the model-predicted noise or injected noise as a guidance signal to steer latents toward target outputs without fine-tuning the model. Builds and evaluates noise-based self-guidance and latent-update rules to control convergence, stability, and final visual or behavioral quality of the generated results.
This work addresses the challenge of achieving efficient, stable, and high-fidelity controllable generation with diffusion models without requiring retraining or backpropagation. By uncovering the intrinsic geometric structure underlying information preservation during the diffusion process, the authors construct a spectral basis derived from the singular functions of the conditional expectation operator. This enables projection of arbitrary guidance signals—such as class labels, CLIP embeddings, or spatial masks—onto the sampling trajectory, yielding training-free, precise control. The method further identifies, for the first time, a phase transition phenomenon within the diffusion process and locates an optimal guidance window, thereby unifying support for multimodal control. On CIFAR-10, it improves conditional accuracy by 37 percentage points over the strongest training-free baseline while accelerating sampling by a factor of four.
This work addresses the challenge of classifier-free guidance in diffusion models—without additional training, access to historical weights, or class conditioning. We propose Sliding Window Guidance (SWG), which leverages the primary model itself as an auxiliary model, constraining its receptive field via a sliding window to enable spatial dependency modeling and alignment of error patterns. Our key insight is that the auxiliary model need not be highly accurate; it suffices for it to share error characteristics with the primary model while exhibiting stronger response sensitivity—yielding substantial gains in generation quality. We further introduce an error-driven guidance mechanism and weight regularization for stable control. SWG incurs zero training overhead, requires no architectural modifications or conditional inputs, and matches state-of-the-art guided methods on quantitative metrics (e.g., FID, LPIPS). Moreover, it achieves superior performance in human visual preference evaluations.
This work addresses the issue that existing guidance methods, such as Classifier-Free Guidance, violate probability conservation under strong guidance, causing generated samples to deviate from the data manifold and resulting in distortion or hallucination. By formulating guidance through the continuity equation, the authors decompose its effect into a divergence term and a score-parallel term, revealing the correspondence between prevailing heuristics and theoretically grounded components. They propose AdaMaG, a plug-in adaptive guidance strategy that incurs no additional inference cost and employs a time-dependent scheduling mechanism to dynamically balance the contributions of these two terms. This approach preserves manifold structure while enabling high-fidelity generation. Experiments demonstrate that AdaMaG significantly enhances realism, suppresses hallucination on image generation benchmarks, and achieves controllable desaturation under high guidance strengths, outperforming current state-of-the-art methods.
Diffusion models’ generation quality and prompt adherence are highly sensitive to the initial Gaussian noise, yet existing noise optimization approaches rely on auxiliary data, additional networks, or backpropagation—limiting practicality. To address this, we propose a lightweight, training-free noise-level guidance framework grounded in forward-process modeling. Our method adaptively optimizes the initial noise during inference by maximizing the likelihood alignment between the noise and generic guidance signals (e.g., CLIP embeddings or classifier-free guidance), enabling end-to-end noise adaptation without modifying the diffusion process. It is agnostic to conditioning mode (conditional or unconditional) and compatible with diverse guidance schemes, introduces no new parameters or computational overhead, and preserves inference efficiency. Evaluated on five standard benchmarks, our approach consistently improves both image fidelity (reducing FID) and text–image alignment (increasing CLIP Score), demonstrating strong generalization and plug-and-play usability.
To address frequent structural artifacts (e.g., distorted hands, faces, arms) and low-fidelity samples in text-to-image generation, this paper proposes a self-guidance mechanism that requires no retraining of diffusion or flow models. The method dynamically identifies and suppresses anomalous generation paths by analyzing the decay characteristics of sampling probability density across noise levels, enabling plug-and-play quality enhancement. Its core innovation is the first training-free guidance paradigm relying solely on intrinsic model sampling probabilities—eliminating dependence on task-specific training data, auxiliary networks, or architectural priors. Compatible with both UNet- and Transformer-based architectures, it supports mainstream models including Stable Diffusion 3.5 and FLUX. Experiments demonstrate superior performance over existing guidance methods in FID and human preference scores, while significantly improving anatomical plausibility and effectively eliminating canonical structural artifacts.
Existing audio generation methods often rely on model retraining or computationally expensive inference-time guidance to achieve fine-grained control, struggling to balance precision and efficiency. This work proposes a low-resource latent-space guidance approach that enables precise, multi-dimensional control over attributes such as intensity, pitch, and tempo directly within the latent variable space of a diffusion model. By integrating a selective Targeted Feature Guidance (TFG) strategy with lightweight Latent Control Heads (LatCHs), the method achieves high-quality audio synthesis with only 7M additional parameters and approximately four hours of training on the Stable Audio Open model. The approach significantly reduces computational overhead while outperforming conventional end-to-end guidance techniques in both controllability and generation quality.
Current classifier-free guidance (CFG) in diffusion models struggles to balance conditional fidelity with generation diversity and lacks effective strategies for scheduling guidance strength. This work proposes an information-theoretic adaptive CFG scheduling framework, introducing information theory into CFG optimization for the first time. By defining a target trade-off through the clean endpoint distribution and dynamically adjusting guidance strength across noise levels without explicit density estimation, the method leverages trajectory-level samples, score evaluations, and information-theoretic metrics to design an adaptive scheduling algorithm. Evaluated on state-of-the-art models such as EDM-XXL and SD-XL, the approach achieves superior distribution alignment. Experiments on ImageNet-512 and COCO demonstrate that the learned scheduling policy significantly enhances generation diversity while preserving conditional consistency, outperforming or matching fixed guidance weights.
This work addresses the instability of existing heuristic post-training guidance methods for generative flows by formulating guidance as a Lyapunov control problem. It establishes, for the first time, an equivalence between guided flow matching and Lyapunov control, and introduces a pseudo-projection operator that unifies diverse guidance strategies—such as classifier-, reward-, or energy-based guidance—under both model-driven and data-driven settings. This framework provides explicit stability guarantees while preserving computational efficiency. Empirical results demonstrate that the proposed approach significantly improves sample quality, guidance fidelity, and robustness across a range of tasks, including synthetic data generation, image inverse problems, reinforcement learning planning, and energy-based modeling.
This work addresses the issue of gradient error accumulation in diffusion models during iterative generation, which arises from misalignment between training objectives and inference dynamics and ultimately impairs generalization. Building upon the “weak-to-strong” guidance principle, the study systematically characterizes the effective operating regimes of classifier-free guidance (CFG) and adaptive guidance (AG) for the first time, and introduces a segmented guidance (SGG) strategy. SGG enhances generation quality by blending guidance signals in distinct segments, jointly optimizing performance during both inference and training. Notably, SGG improves inference without requiring additional training and integrates seamlessly into architectures such as Stable Diffusion 3/3.5 and Transformers. Experiments demonstrate that SGG significantly boosts model generalization in both conditional and unconditional generation tasks, outperforming existing training-free guidance methods.
Existing generative models often suffer from inefficiency when guided by user-specified rewards—such as aesthetic quality or human preferences—due to computationally expensive procedures or multi-step approximations. This work reframes the guidance problem as a deterministic optimal control task and, for the first time, naturally integrates flow matching into its solution framework, yielding a training-free, single-trajectory guidance method. Requiring only three function evaluations (NFEs), the proposed approach achieves high-quality alignment in text-to-image generation and matches or surpasses state-of-the-art methods across diverse settings, including inverse problems, style transfer, human preference optimization, and VLM-based rewards. It accelerates inference by over an order of magnitude while providing a unified theoretical foundation that subsumes existing guidance algorithms.