🤖 AI Summary
This study addresses the inherent challenge of balancing content leakage prevention and style fidelity in diffusion-based style transfer by proposing a training-free collaborative framework. The core methodology identifies the leakage-degradation dilemma across the entire generation pipeline and establishes a multi-vision foundation model collaboration mechanism, theoretically proving that anchor error decreases as the number of models increases. Furthermore, it introduces orthogonal subspace projection, ensemble inversion, and energy-guided calibration techniques to suppress leakage throughout the generation process while optimizing sampling. Experimental evaluations on the StyleBench benchmark demonstrate that the proposed approach significantly reduces content leakage and substantially improves both style alignment and LLM-based evaluation scores.
📝 Abstract
Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{https://github.com/0606zt/CLeaR}{https://github.com/0606zt/CLeaR}.