TemplateCraft: Agentic Visual Template Generation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the time-consuming nature of manual orchestration in short video template generation by proposing a Qwen3-VL-based multi-agent collaborative framework that automatically translates natural language instructions into executable templates. Methodologically, it designs a pipeline encompassing planning, asset and effect workflow generation, and protocol compilation. A novel Planner-Evaluator closed-loop mechanism is introduced, leveraging execution feedback for targeted rollback and integrating stage-level with long-term memory to enable parameter-free iterative revision, complemented by persistent asset storage for enhanced reusability. Evaluated on the TemplateBench benchmark, the framework increases success rates for image and video template generation to 66.7% and 50.0%, respectively, significantly improving template compliance and stylistic consistency.
📝 Abstract
The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset preparation and tool orchestration. We propose TemplateCraft, a multi-agent system that converts natural-language instructions into client-executable templates through planning, material generation, effect-workflow generation, and protocol compilation. Its Planner-Evaluator loop uses execution feedback for targeted rollback, while stage-level and long-term memory support revision without parameter updates. We evaluate TemplateCraft on TemplateBench, derived from 60 real-world templates. With the same Qwen3-VL backbone, TemplateCraft raises image/video generation success rates from 56.7%/30.0% to 66.7%/50.0% over Planner-only (best-of-three) and improves template adherence and style consistency. With additional evaluation and revision, it matches or exceeds a GPT-4o Planner-only baseline on selected metrics. Persistent assets further improve cross-input style consistency.
Problem

Research questions and friction points this paper is trying to address.

visual template generation
short video creation
one-click content creation
template automation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent System
Visual Template Generation
Planner-Evaluator Loop
Memory Mechanism
Execution Feedback
💼 Related Jobs
No related jobs found.
Hongjie Yu
Hongjie Yu
Fudan University
Zhiyuan Fan
Zhiyuan Fan
PhD Student, MIT
reinforcement learningcomputational game theory
Y
Yuzhe Zhang
Peking University
J
Jiangcun Du
Kuaishou Technology
Z
Zhicheng Gao
Kuaishou Technology
Y
Yuhong Zhang
Kuaishou Technology
X
Xiaokai Zhan
Kuaishou Technology
Z
Zongshi Xie
Kuaishou Technology