🤖 AI Summary
This work addresses the high computational cost of multi-step diffusion models commonly used in identity-preserving image generation. The authors propose a training-free acceleration strategy that directly transfers a pre-trained InfuseNet identity adapter to a distilled schnell backbone network, replacing only the backbone and disabling classifier-free guidance. Through attention flow norm probing and evaluation with ArcFace and LPIPS metrics, they demonstrate that identity fidelity reaches an effective range within 4–8 denoising steps, with subsequent steps primarily refining visual details. Compared to the 28-step baseline, the proposed method reduces inference latency by 5.9× while improving ArcFace similarity by 0.028 and lowering LPIPS by 0.016, achieving a superior trade-off between efficiency and identity preservation.
📝 Abstract
Identity-preserved image generation is typically built on many-step diffusion backbones, making personalized generation expensive at deployment time. We show that this cost is often unnecessary for identity-conditioned FLUX generation. A frozen InfuseNet identity adapter trained with dev transfers directly to the distilled schnell backbone without retraining. This two-line replacement -- changing the backbone path and disabling classifier-free guidance -- reduces latency by 5.9x while improving ArcFace identity similarity by +0.028 and lpips by -0.016 over the standard 28-step dev baseline. To explain why this works, we analyze the denoising trajectory and find that identity fidelity enters an early effective regime, often within 4-8 steps, while later steps primarily refine visual detail, sharpness, and contrast. Adapter ablations confirm that identity formation depends on the identity adapter, while attention-stream norm probes suggest that the relative conditioning contribution decreases as sampling proceeds. Preliminary style-adapter and object-adapter sweeps on SDXL and SD1.5 show similar diminishing returns after intermediate steps. These results position distilled backbone replacement as a simple, training-free strategy for improving the efficiency-fidelity tradeoff of identity-preserved generation.