Score
Designs algorithms and modules that plan, control, or impose priors on the sequence of latent states or noise-to-signal transitions in diffusion-based generative processes — including trajectory planners, editing trajectories, flow-based controllers, and learned generative priors over trajectories. Builds and analyzes methods to steer semantic manipulations across timesteps, preserve consistency of edits, and reduce unintended identity or attribute retention during sampling and editing.
To address the lack of goal-directedness in diffusion model inference, this paper proposes a fine-tuning-free unified guidance framework for real-time optimization of task-critical properties—such as structural stability and binding affinity—in downstream applications like protein design. Methodologically, it first unifies sequence Monte Carlo (SMC) guidance, value-function-based sampling, and classifier guidance under a coherent soft-optimal denoising paradigm. It further integrates reinforcement learning policy approximation, Monte Carlo tree search, and cross-modal alignment to enhance target alignment during sampling. Experiments demonstrate substantial improvements in biologically relevant metrics of generated proteins, with an open-source codebase ensuring reproducibility. The core contribution is a general, efficient, training-free diffusion sampling guidance paradigm that bridges generative modeling and structured biological design objectives.
This study addresses the computational bottleneck associated with dynamically constrained sampling of state spaces in feedback control and planning. To overcome this challenge, the work reformulates control as a dynamically constrained sampling problem, establishing a mapping framework that bridges controllability, optimal control theory, and generative modeling. Specifically, it integrates flow matching, normalizing flows, and denoising diffusion techniques to guide system evolution toward target states or distributions. The proposed approach enables efficient reachable set sampling and precise trajectory planning while unifying control-theoretic and generative-modeling paradigms. Furthermore, the authors provide an accessible open-source tutorial to facilitate practical adoption by the research community.
This work proposes Chain-of-Trajectories (CoTj), a training-free framework that introduces System 2–style deliberative planning into diffusion sampling to overcome the inefficiency of fixed, content-agnostic scheduling in navigating high-dimensional noise manifolds. CoTj extracts low-dimensional semantic features via Diffusion DNA and reformulates the sampling process as a path-planning problem on a directed acyclic graph. By adopting a predict–plan–execute paradigm, it enables context-aware dynamic allocation of computational resources. Experiments demonstrate that CoTj consistently enhances generation quality and stability across diverse diffusion models while reducing redundant computations, thereby achieving efficient, resource-aware sampling.
This study addresses the challenge of simultaneously maintaining trajectory feasibility and consistency during test-time adaptation of diffusion models by proposing a persistent particle planning framework. The method unifies denoising and replanning as a Feynman–Kac particle system, optimizing control decisions through a maintained population of candidate plans combined with partial renoising and constraint-aware sampling. We theoretically demonstrate that preserving rare paths is more efficient and stable than sampling from scratch. Experiments show that the proposed framework effectively reduces route switching and constraint violations while decreasing the required number of iterations, achieving faster convergence and higher success rates in maze navigation tasks.
Discrete-state-space generative models—e.g., for small molecules, DNA, and protein sequences—lack principled, controllable guidance mechanisms. Existing continuous-domain guidance paradigms do not generalize to discrete spaces, hindering attribute-controlled generation. Method: This paper introduces the first general, differentiable guidance framework for discrete generative models based on continuous-time Markov processes, unifying discrete diffusion and flow-matching architectures. It overcomes the fundamental incompatibility of continuous guidance with discrete state spaces by establishing a theoretically grounded guidance theory for discrete domains. The framework enables arbitrary differentiable guidance objectives without model retraining, leveraging probability path reweighting and gradient-driven discrete sampling. Contribution/Results: Experiments demonstrate substantial improvements in target property satisfaction rates and sample diversity across diverse biomolecular generation tasks, while maintaining flexibility and strong generalization across guidance objectives and model architectures.
Existing reward-based fine-tuning of generative models lacks theoretical foundations—particularly within flow matching and denoising diffusion frameworks—struggling to jointly ensure accuracy, sample diversity, and generalization to unseen human preferences. Method: This work formally recasts reward-driven generation as a stochastic optimal control (SOC) problem. We rigorously prove that memoryless noise scheduling is necessary to decouple noise from samples, thereby enabling tractable optimization. Building on this insight, we propose Adjoint Matching: a novel algorithm that transforms the SOC formulation into a supervised regression task via adjoint-state methods, unifying stochastic optimal control, adjoint calculus, flow matching, and diffusion modeling. Results: Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in fidelity, perceptual realism, and generalization to unseen reward models, while preserving high sample diversity.
This work addresses the challenge that diffusion-based policies often fail to discover rare yet effective behavioral patterns under scarce demonstration data, frequently converging to suboptimal solutions or generating infeasible trajectories. To overcome this limitation, the authors propose a novel framework integrating a Feynman–Kac corrector with a learnable guidance potential, which systematically steers the diffusion process toward under-explored feasible regions of the trajectory space. By coupling this guided exploration with sample-based trajectory optimization and an iterative retraining mechanism, the method continuously refines and reuses newly discovered behaviors. Empirical results demonstrate that the approach substantially enhances policy diversity, feasibility, and generalization, consistently uncovering novel and effective strategies across diverse manipulation tasks—surpassing the capabilities of conventional sampling and reinforcement learning methods.
Existing generative models typically optimize only the final output, lacking semantic controllability over intermediate generation steps. This work proposes Trajectory Forcing, a novel framework that explicitly models the generation path as a semantically editable trajectory progressing from global layout to fine-grained details. The approach constructs a coarse-to-fine teacher hierarchy based on DINOv2 clustering and trains a hierarchical, step-conditioned flow-matching model to enable structured and editable intermediate state generation. Experiments demonstrate that the framework maintains high-fidelity image synthesis while supporting cross-scale local editing. Furthermore, it introduces trajectory-aware metrics for consistency and controllability, moving beyond conventional endpoint-only evaluations such as FID and thereby addressing their inherent limitations.
This work proposes a unified framework that integrates safety constraints and task objectives into pretrained diffusion models and flow-matching policies during inference without requiring retraining. By formulating constrained trajectory generation as a Bayesian posterior sampling problem—using expert demonstration distributions as the prior and cost functions to define the likelihood—the method extends the Feynman–Kac corrector, for the first time, to deterministic flow-matching strategies, enabling consistent correction across both classes of generative models. Evaluated on Diffusion Policy, GR00T-N1.6, and π₀.₅, the approach demonstrates zero-shot capability in handling complex, non-convex obstacle avoidance tasks and significantly outperforms the original π₀.₅ policy in terms of constraint satisfaction and task performance.
This work addresses the challenge of achieving efficient, stable, and high-fidelity controllable generation with diffusion models without requiring retraining or backpropagation. By uncovering the intrinsic geometric structure underlying information preservation during the diffusion process, the authors construct a spectral basis derived from the singular functions of the conditional expectation operator. This enables projection of arbitrary guidance signals—such as class labels, CLIP embeddings, or spatial masks—onto the sampling trajectory, yielding training-free, precise control. The method further identifies, for the first time, a phase transition phenomenon within the diffusion process and locates an optimal guidance window, thereby unifying support for multimodal control. On CIFAR-10, it improves conditional accuracy by 37 percentage points over the strongest training-free baseline while accelerating sampling by a factor of four.
This study addresses the challenge in multi-robot planning of reconciling the incorporation of coarse-grained trajectory priors with the preservation of solution multimodality. To this end, it proposes a timestep-dependent, trajectory-level guidance mechanism based on diffusion models. Specifically, priors are progressively injected into the clean trajectory space during reverse generation, with guidance strength dynamically adjusted, while gradient descent is employed to jointly optimize planning costs and collision constraints. This approach enables controllable multi-agent trajectory synthesis, ensuring collision avoidance safety while generating diverse and feasible multimodal planning solutions.