Score
Design and implement finetuning procedures for generative models that calibrate kernel-based distributional discrepancies (e.g., MMD) by choosing and tuning kernels, feature mappings, and optimization schedules. Build loss functions and regularizers that minimize the MMD between generated and target feature distributions while constraining updates toward the pretrained model to preserve sample validity, diversity, and other generation constraints.
This work addresses the common discrepancy between existing generative models and target data in critical feature distributions, a problem exacerbated by direct fine-tuning that often leads to overfitting and poor controllability. To overcome this, the authors propose kCGM, a method that calibrates diverse pre-trained generative models—including autoregressive, continuous, and discrete diffusion models—using only feature-level supervision. kCGM minimizes the maximum mean discrepancy (MMD) in feature space between generated samples and target data while incorporating KL divergence regularization. The approach effectively aligns feature distributions without compromising generation quality. Experiments demonstrate that, in antibiotic molecule generation, kCGM substantially outperforms direct fine-tuning in both chemical validity and alignment with desired features, and it generalizes successfully to protein and DNA sequence generation tasks.
Generative models often suffer from poor calibration, where predicted class probabilities diverge from empirical sampling statistics. This paper proposes a general constraint-based calibration framework: it minimizes the KL divergence between the calibrated and original model distributions, subject to multiple category- or statistic-specific constraints. Two scalable optimization objectives are introduced: (i) relaxed loss—incorporating calibration error as a regularization term—and (ii) reward loss—transforming constraints into differentiable reward signals for fine-tuning. Both support joint optimization over hundreds of constraints. The method preserves generation quality while substantially reducing calibration error—even for billion-parameter models. Extensive experiments across protein design, image generation, and language modeling demonstrate strong generalization and computational efficiency.
This work addresses the challenge of distribution shift in pre-trained diffusion models under domain adaptation scenarios, where limited target reference samples and the inability to retrain often lead to biased generation. The authors propose a novel inference-time approach that guides the reverse diffusion process using gradients derived from Maximum Mean Discrepancy (MMD). For the first time, MMD is employed as a differentiable, low-variance distributional distance metric directly within diffusion guidance, enabling precise alignment between the generated distribution and a given target reference set without any retraining. The method naturally supports prompt-aware adaptation in conditional generation and extends efficiently to latent diffusion models (LDMs). Experiments on both synthetic and real-world benchmarks demonstrate its effectiveness in achieving distribution alignment while preserving high sample quality and fidelity.
Traditional implicit generative models (e.g., GANs) suffer from training instability, mode collapse, and inaccurate tail characterization when modeling heavy-tailed, high-dimensional multivariate distributions. To address these challenges, this paper proposes Pareto-ISL—a novel implicit score learning framework. Its core contributions are: (1) the first integration of generalized Pareto noise into implicit score learning (ISL), explicitly capturing heavy-tailed behavior; and (2) a random-projection-based multidimensional ISL loss, extending ISL beyond univariate settings to scalable high-dimensional implicit modeling. Experiments demonstrate that Pareto-ISL accurately reproduces both central and tail regions of multivariate heavy-tailed distributions, significantly mitigates mode collapse, exhibits robustness to hyperparameter choices, and scales linearly in computational complexity with dimensionality.
This study investigates how pretraining model scale, pretraining dataset size, and training strategies affect out-of-distribution (OOD) generalization and confidence calibration. We conduct controlled experiments across 100 models on four OOD benchmarks—ImageNet-C, -R, -A, and -O—accumulating over 120,000 GPU hours. Results reveal that pretraining model selection alone substantially improves OOD accuracy and calibration (reducing Expected Calibration Error by up to 40%), outperforming most dedicated OOD algorithms. Contrary to the prevailing belief that larger models degrade calibration, we find that scaling both model capacity and pretraining data jointly enhances calibration. Moreover, overconfidence systematically diminishes with increased model and data scale. This work provides the first large-scale empirical evidence demonstrating a dual positive effect of pretraining configuration—model architecture and data scale—on OOD robustness, establishing critical empirical foundations for designing trustworthy vision models.
This work addresses the challenge that generative models often fail to accurately recover the true data distribution when modeling low-dimensional manifolds subject to equality constraints, primarily due to restricted support sets. To overcome this limitation, the authors propose a constraint-aware distribution perturbation method that extends the distribution’s support into the ambient space while preserving the underlying manifold geometry. This approach is compatible with mainstream generative frameworks such as diffusion models and normalizing flows, offering computational efficiency, mathematical rigor, and flexibility. It effectively mitigates sampling instability and distributional distortion commonly observed in conventional constrained modeling. Experimental results demonstrate that the proposed method significantly outperforms existing approaches across multiple scientific data tasks, consistently generating samples that faithfully reproduce the original data distribution.
This work addresses the lack of theoretical guarantees for generalization in existing discriminator-guided generative models. Building upon the strong duality of f-divergences, we propose a universal discriminator-guided refinement framework that enhances the generalization capability of any generative model, including diffusion models. We provide the first theoretical proof that this guidance mechanism provably reduces the generalization gap and establish a quantitative relationship between the gap reduction and the Rademacher complexity of the discriminator class. Furthermore, our framework offers a unified theoretical explanation for the empirical success of recent score-based diffusion methods. While maintaining broad applicability, the proposed approach delivers rigorous theoretical justification for widely used yet previously heuristic refinement strategies.
Existing diffusion models often suffer from low constraint satisfaction rates or degraded sample quality when generating samples under complex feasibility constraints, primarily due to distributional mismatches between training and sampling phases. This work proposes a trajectory-aware fine-tuning framework based on online rollouts, which integrates constraint guidance during training and, for the first time, incorporates a numerical integration perspective into the diffusion process. By end-to-end differentiating denoising trajectories under a fixed noise schedule, the method explicitly exposes constraint violations, thereby aligning the training and sampling distributions. Combining constraint-aware guidance, differentiable noise scheduling, and an online rollout mechanism, the approach significantly improves constraint satisfaction across multiple tasks while maintaining generation quality on par with current state-of-the-art methods.
We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is $\pi\propto\mu e^{\tau r}$, where $r$ is the reward, $\tau>0$ the inverse temperature, and $\mu$ denotes the pretrained model's terminal density for fine-tuning or the constant $1$ for sampling. We shift the paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities. Under compatible smooth-realization assumptions, canonical velocities form a manifold diffeomorphic to the density manifold. Transporting the Fisher-Rao metric and mixture connection to this manifold, we show that the reverse-KL Hessian equals the metric, so the Newton direction coincides with the negative Fisher-Rao gradient. At terminal density $\rho$, each stage takes a tangential step generated by the regularized reward $r-\frac1\tau\log(\rho/\mu)$, followed by terminal-density-preserving canonicalization. This canonical retraction yields an exact finite-stepsize density characterization. For the ideal iteration, we prove strict reverse-KL descent away from the target for $0<\eta \le \tau$, global convergence under mild conditions, and local quadratic convergence for full steps ($\eta=\tau$). Covariance and gradient forms, each with forward or reverse regression-pair constructions, yield sample-wise tangential-update losses with the same population minimizer, without importance sampling or full-trajectory backpropagation. We develop approximate updates and define critical-point consistency as vanishing tangential displacement if and only if $\rho=\pi$. We recover representative methods as exact realizations, critical-point-consistent approximations, or objective-altering variants, enabling modular algorithm design. Our work advances the theory and algorithms of reinforcement learning for generative models.
This work addresses a critical limitation in existing fairness definitions for generative models, which focus solely on balancing generation probabilities across sensitive groups while neglecting disparities in generation quality, thereby yielding fragile fairness evaluations. To remedy this, we propose a novel paradigm termed Equal Generative Treatment (EGT), which requires that all sensitive groups exhibit comparable generation quality as measured by f-divergence, and we uncover an intrinsic coupling between EGT and overall model performance. To operationalize this principle, we develop a min-max optimization–based fine-tuning strategy that significantly improves fairness in generation quality across groups in both image and text generation tasks, without compromising overall model competitiveness. This study is the first to formally incorporate generation quality into the fairness framework, establishing a theoretical foundation for fair generative modeling.