Calibrating Generative Models to Feature Distributions with MMD Finetuning

📅 2026-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the common discrepancy between existing generative models and target data in critical feature distributions, a problem exacerbated by direct fine-tuning that often leads to overfitting and poor controllability. To overcome this, the authors propose kCGM, a method that calibrates diverse pre-trained generative models—including autoregressive, continuous, and discrete diffusion models—using only feature-level supervision. kCGM minimizes the maximum mean discrepancy (MMD) in feature space between generated samples and target data while incorporating KL divergence regularization. The approach effectively aligns feature distributions without compromising generation quality. Experiments demonstrate that, in antibiotic molecule generation, kCGM substantially outperforms direct fine-tuning in both chemical validity and alignment with desired features, and it generalizes successfully to protein and DNA sequence generation tasks.
📝 Abstract
Generative models can produce individually plausible samples while deviating substantially from a target set in the distribution of key features. For example, a model pretrained on broad drug-like chemical space may generate molecules whose molecular features differ from those of a therapeutic class of interest, such as known antibiotics. Correcting such distributional miscalibration is challenging: direct finetuning on the target set can overfit and does not control which features are matched. To fill this gap, we introduce kernel Calibrating Generative Models (kCGM). kCGM minimizes a maximum mean discrepancy (MMD) between generated and target feature distributions using an unbiased score-function estimator, with KL regularization to remain close to the pretrained model. On a target set of 174 antibiotics, direct finetuning sacrifices chemical validity for feature-distribution matching, whereas kCGM improves target feature matching while increasing validity. We further demonstrate kCGM in protein and DNA generation tasks, showing it can adapt autoregressive, continuous-space diffusion, and discrete diffusion models using only feature-level supervision. Code is available at https://github.com/smithhenryd/cgm.
Problem

Research questions and friction points this paper is trying to address.

generative models
distribution calibration
feature distribution
maximum mean discrepancy
molecular generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

maximum mean discrepancy
generative model calibration
feature distribution matching
score-function estimator
KL regularization
🔎 Similar Papers
No similar papers found.
N
Nathaniel L. Diamant
Stanford University
B
Brian L. Trippe
Stanford University