Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation that adapting few-step distilled text-to-image models to new styles typically requires additional training, which compromises their native few-step generation capabilities. To overcome this, we propose StyleForge, a training-free framework that employs visual evidence organization to automatically extract global rendering and local illumination features from reference images. These features are compiled into reusable, unified textual instructions that elicit the model’s intrinsic style transfer ability with frozen weights. As the first fully automated, training-free style representation mechanism, it enables efficient cross-content reuse. Experimental results demonstrate that StyleForge improves generation quality by up to 29.47% over the strongest baseline, achieving strong stylization while preserving high content consistency.
πŸ“ Abstract
Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit stylistic capabilities through language. Pursuing this direction requires textual guidance that captures how visual attributes jointly define a style and remain applicable as the depicted content changes. To explore this approach, we introduce StyleForge, a fully automatic, training-free framework that expresses reference styles as reusable rendering instructions. By integrating overall rendering characteristics with local color and lighting behavior, StyleForge organizes visual evidence from reference images into a coherent specification of how the target style should be expressed. The specification is then compiled into textual guidance that can be reused across content prompts, enabling frozen step-distilled T2I models to render different subjects and scenes in the reference style while retaining native few-step generation. Extensive experiments show relative gains of up to 29.47\% in generation quality scores over the strongest baseline, while Pareto analysis indicates that improved stylization is accompanied by strong adherence to the requested content.
Problem

Research questions and friction points this paper is trying to address.

step-distilled diffusion models
style transfer
training-free
text-to-image generation
textual guidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Step-Distilled Diffusion Models
Training-Free Style Transfer
Textual Guidance
Rendering Instructions
StyleForge
πŸ”Ž Similar Papers