SPAST: Arbitrary style transfer with style priors via pre-trained large-scale model

📅 2025-05-01
🏛️ Neural Networks
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing arbitrary style transfer methods face a fundamental trade-off: lightweight models yield low-fidelity outputs with prominent artifacts, whereas large models achieve higher visual quality but suffer from poor content-structure preservation and slow inference. This paper proposes a fine-tuning-free lightweight framework that, for the first time, embeds learnable explicit style priors into the frozen CLIP and diffusion Transformer (DiT) feature spaces, enabling effective content-style disentanglement. Our approach integrates CLIP-aligned guidance, DiT feature distillation, plug-and-play style adapters, and a contrastive style reconstruction loss—enabling zero-shot generalization to unseen styles and cross-domain reuse. On MSCOCO→WikiArt, our method reduces FID by 37%, improves style similarity by 2.1×, and achieves 48 fps inference speed at 512×512 resolution—significantly outperforming AdaIN, StyleCLIP, and LDM-Stylize.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: Sentiment Analysis, Stylistic Analysis, and Argument Mining

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
Problem

Research questions and friction points this paper is trying to address.

Improving quality of stylized images in style transfer
Reducing inference time in large-scale model-based methods
Preserving content structure while applying style features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Local-global Window Size Stylization Module
Incorporates style prior loss from large model
Reduces inference time while preserving quality
🔎 Similar Papers
2024-07-01arXiv.orgCitations: 3
Zhanjie Zhang
Zhanjie Zhang
Zhejiang University
computer vision
Q
Quanwei Zhang
College of Computer Science and Technology Zhejiang University No. 38 Zheda Road Hangzhou 310000 China
Junsheng Luan
Junsheng Luan
Zhejiang University
Mengyuan Yang
Mengyuan Yang
Zhejiang University
Y
Yun Wang
Department of Computer Science City University of Hong Kong Tat Chee Avenue Kowloon Hong Kong SAR
L
Lei Zhao
College of Computer Science and Technology Zhejiang University No. 38 Zheda Road Hangzhou 310000 China