mmd finetuning

Design and implement finetuning procedures for generative models that calibrate kernel-based distributional discrepancies (e.g., MMD) by choosing and tuning kernels, feature mappings, and optimization schedules. Build loss functions and regularizers that minimize the MMD between generated and target feature distributions while constraining updates toward the pretrained model to preserve sample validity, diversity, and other generation constraints.

mmdfinetuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the common discrepancy between existing generative models and target data in critical feature distributions, a problem exacerbated by direct fine-tuning that often leads to overfitting and poor controllability. To overcome this, the authors propose kCGM, a method that calibrates diverse pre-trained generative models—including autoregressive, continuous, and discrete diffusion models—using only feature-level supervision. kCGM minimizes the maximum mean discrepancy (MMD) in feature space between generated samples and target data while incorporating KL divergence regularization. The approach effectively aligns feature distributions without compromising generation quality. Experiments demonstrate that, in antibiotic molecule generation, kCGM substantially outperforms direct fine-tuning in both chemical validity and alignment with desired features, and it generalizes successfully to protein and DNA sequence generation tasks.

distribution calibrationfeature distributiongenerative models

Calibrating Generative Models

Oct 11, 2025
HD
Henry D. Smith
🏛️ Stanford University

Generative models often suffer from poor calibration, where predicted class probabilities diverge from empirical sampling statistics. This paper proposes a general constraint-based calibration framework: it minimizes the KL divergence between the calibrated and original model distributions, subject to multiple category- or statistic-specific constraints. Two scalable optimization objectives are introduced: (i) relaxed loss—incorporating calibration error as a regularization term—and (ii) reward loss—transforming constraints into differentiable reward signals for fine-tuning. Both support joint optimization over hundreds of constraints. The method preserves generation quality while substantially reducing calibration error—even for billion-parameter models. Extensive experiments across protein design, image generation, and language modeling demonstrate strong generalization and computational efficiency.

Calibration is framed as constrained optimization with KL divergenceGenerative models often have miscalibrated class probabilitiesMethods reduce calibration errors across billion-parameter models

This work addresses the challenge of distribution shift in pre-trained diffusion models under domain adaptation scenarios, where limited target reference samples and the inability to retrain often lead to biased generation. The authors propose a novel inference-time approach that guides the reverse diffusion process using gradients derived from Maximum Mean Discrepancy (MMD). For the first time, MMD is employed as a differentiable, low-variance distributional distance metric directly within diffusion guidance, enabling precise alignment between the generated distribution and a given target reference set without any retraining. The method naturally supports prompt-aware adaptation in conditional generation and extends efficiently to latent diffusion models (LDMs). Experiments on both synthetic and real-world benchmarks demonstrate its effectiveness in achieving distribution alignment while preserving high sample quality and fidelity.

diffusion modelsdistribution adaptationdomain mismatch

Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss

Oct 29, 2024
JM
José Manuel de Frutos
🏛️ Universidad Carlos III de Madrid | Instituto de Investigación Sanitaria Gregorio Marañón (IiSGM)

Traditional implicit generative models (e.g., GANs) suffer from training instability, mode collapse, and inaccurate tail characterization when modeling heavy-tailed, high-dimensional multivariate distributions. To address these challenges, this paper proposes Pareto-ISL—a novel implicit score learning framework. Its core contributions are: (1) the first integration of generalized Pareto noise into implicit score learning (ISL), explicitly capturing heavy-tailed behavior; and (2) a random-projection-based multidimensional ISL loss, extending ISL beyond univariate settings to scalable high-dimensional implicit modeling. Experiments demonstrate that Pareto-ISL accurately reproduces both central and tail regions of multivariate heavy-tailed distributions, significantly mitigates mode collapse, exhibits robustness to hyperparameter choices, and scales linearly in computational complexity with dimensionality.

Addresses unstable training and mode dropping in traditional generative modelsExtends invariant statistical loss to handle heavy-tailed multivariate distributionsOvercomes computational challenges of high-dimensional data with random projections

An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration

Jul 17, 2023
HN
Hiroki Naganuma
🏛️ Mila | Université de Montréal | RIKEN AIP

This study investigates how pretraining model scale, pretraining dataset size, and training strategies affect out-of-distribution (OOD) generalization and confidence calibration. We conduct controlled experiments across 100 models on four OOD benchmarks—ImageNet-C, -R, -A, and -O—accumulating over 120,000 GPU hours. Results reveal that pretraining model selection alone substantially improves OOD accuracy and calibration (reducing Expected Calibration Error by up to 40%), outperforming most dedicated OOD algorithms. Contrary to the prevailing belief that larger models degrade calibration, we find that scaling both model capacity and pretraining data jointly enhances calibration. Moreover, overconfidence systematically diminishes with increased model and data scale. This work provides the first large-scale empirical evidence demonstrating a dual positive effect of pretraining configuration—model architecture and data scale—on OOD robustness, establishing critical empirical foundations for designing trustworthy vision models.

Effect of pre-training dataset size on OOD performance and confidenceImpact of pre-trained model size on OOD generalization and calibrationInfluence of training strategies on model accuracy and overconfidence mitigation

Latest Papers

What's happening recently
View more

This work addresses the challenge that generative models often fail to accurately recover the true data distribution when modeling low-dimensional manifolds subject to equality constraints, primarily due to restricted support sets. To overcome this limitation, the authors propose a constraint-aware distribution perturbation method that extends the distribution’s support into the ambient space while preserving the underlying manifold geometry. This approach is compatible with mainstream generative frameworks such as diffusion models and normalizing flows, offering computational efficiency, mathematical rigor, and flexibility. It effectively mitigates sampling instability and distributional distortion commonly observed in conventional constrained modeling. Experimental results demonstrate that the proposed method significantly outperforms existing approaches across multiple scientific data tasks, consistently generating samples that faithfully reproduce the original data distribution.

constrained samplingdistribution recoveryequality constraints

This work addresses the lack of theoretical guarantees for generalization in existing discriminator-guided generative models. Building upon the strong duality of f-divergences, we propose a universal discriminator-guided refinement framework that enhances the generalization capability of any generative model, including diffusion models. We provide the first theoretical proof that this guidance mechanism provably reduces the generalization gap and establish a quantitative relationship between the gap reduction and the Rademacher complexity of the discriminator class. Furthermore, our framework offers a unified theoretical explanation for the empirical success of recent score-based diffusion methods. While maintaining broad applicability, the proposed approach delivers rigorous theoretical justification for widely used yet previously heuristic refinement strategies.

discriminator guidancef-divergencesgeneralization

Existing diffusion models often suffer from low constraint satisfaction rates or degraded sample quality when generating samples under complex feasibility constraints, primarily due to distributional mismatches between training and sampling phases. This work proposes a trajectory-aware fine-tuning framework based on online rollouts, which integrates constraint guidance during training and, for the first time, incorporates a numerical integration perspective into the diffusion process. By end-to-end differentiating denoising trajectories under a fixed noise schedule, the method explicitly exposes constraint violations, thereby aligning the training and sampling distributions. Combining constraint-aware guidance, differentiable noise scheduling, and an online rollout mechanism, the approach significantly improves constraint satisfaction across multiple tasks while maintaining generation quality on par with current state-of-the-art methods.

constrained generationdiffusion modelsdistribution shift

We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is $\pi\propto\mu e^{\tau r}$, where $r$ is the reward, $\tau>0$ the inverse temperature, and $\mu$ denotes the pretrained model's terminal density for fine-tuning or the constant $1$ for sampling. We shift the paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities. Under compatible smooth-realization assumptions, canonical velocities form a manifold diffeomorphic to the density manifold. Transporting the Fisher-Rao metric and mixture connection to this manifold, we show that the reverse-KL Hessian equals the metric, so the Newton direction coincides with the negative Fisher-Rao gradient. At terminal density $\rho$, each stage takes a tangential step generated by the regularized reward $r-\frac1\tau\log(\rho/\mu)$, followed by terminal-density-preserving canonicalization. This canonical retraction yields an exact finite-stepsize density characterization. For the ideal iteration, we prove strict reverse-KL descent away from the target for $0<\eta \le \tau$, global convergence under mild conditions, and local quadratic convergence for full steps ($\eta=\tau$). Covariance and gradient forms, each with forward or reverse regression-pair constructions, yield sample-wise tangential-update losses with the same population minimizer, without importance sampling or full-trajectory backpropagation. We develop approximate updates and define critical-point consistency as vanishing tangential displacement if and only if $\rho=\pi$. We recover representative methods as exact realizations, critical-point-consistent approximations, or objective-altering variants, enabling modular algorithm design. Our work advances the theory and algorithms of reinforcement learning for generative models.

fine-tuninggenerative modelingNewton Matching

This work addresses a critical limitation in existing fairness definitions for generative models, which focus solely on balancing generation probabilities across sensitive groups while neglecting disparities in generation quality, thereby yielding fragile fairness evaluations. To remedy this, we propose a novel paradigm termed Equal Generative Treatment (EGT), which requires that all sensitive groups exhibit comparable generation quality as measured by f-divergence, and we uncover an intrinsic coupling between EGT and overall model performance. To operationalize this principle, we develop a min-max optimization–based fine-tuning strategy that significantly improves fairness in generation quality across groups in both image and text generation tasks, without compromising overall model competitiveness. This study is the first to formally incorporate generation quality into the fairness framework, establishing a theoretical foundation for fair generative modeling.

f-divergencesfairnessgeneration quality

Hot Scholars

KC

Kyung Chul Lee

Yonsei University, Seoul National University
BiophotonicsComputational Imaging
MR

Maximilian Rokuss

German Cancer Research Center (DKFZ), University of Heidelberg
Computer VisionDeep LearningMedical Image Computing
YK

Yannick Kirchhoff

PhD Student, DKFZ
Computer VisionDeep LearningMedical Image Computing
TW

Tassilo Wald

PhD Student, Deutsche Krebsforschungszentrum (DKFZ)
representation learningself-supervised learningmedical image analysis