ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation in subject fidelity observed when fine-tuning diffusion models, such as DreamBooth, on synthetic images. By identifying classifier-free guidance (CFG) as the root cause of this deterioration, we propose a training-free, real-image-free adaptive correction method applied at sampling time in the frequency domain. This approach effectively restores subject details and color balance through frequency band scaling. Experiments conducted across the Stable Diffusion model family demonstrate that the proposed method closes the fidelity gap by 51%–64%, substantially improving subject reconstruction quality while strictly preserving text alignment capabilities.
📝 Abstract
Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.
Problem

Research questions and friction points this paper is trying to address.

subject fidelity
synthetic images
text-to-image diffusion models
personalization
classifier-free guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Classifier-Free Guidance
Training-Free Correction
Subject Fidelity
Frequency Band Scaling
Synthetic Image Personalization
🔎 Similar Papers
No similar papers found.