🤖 AI Summary
This study addresses the absence of large-scale dedicated datasets and systematic evaluation frameworks in product-centric advertising video generation. To this end, we introduce AdSpark, a unified dataset comprising 300,000 samples, alongside AdSpark-Bench, a six-dimensional diagnostic benchmark. Methodologically, we employ a hybrid data collection strategy combining real-world and synthetic sources, providing structured creative annotations and an automated evaluation system encompassing multiple dimensions such as visual quality and product fidelity. Through model fine-tuning experiments, we validate the effectiveness of the proposed dataset and reveal critical bottlenecks in existing models regarding product identity consistency and multi-shot narrative generation. Ultimately, this work bridges the evaluation gap in this specific domain and establishes a new research standard for product-oriented advertising video synthesis.
📝 Abstract
Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce \textbf{AdSpark}, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. \textit{AdSpark-300K} contains approximately 300K reference image--prompt--video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose \textit{AdSpark-Bench}, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.