๐ค AI Summary
To address the challenges of generic ad content, poor landing-page alignment, and high manual review costs in off-site e-commerce marketing, this paper proposes MarketingFMโa retrieval-augmented generation (RAG)-based system for large-scale, personalized ad copy generation. We further introduce AutoEval, an automated evaluation framework integrating rule-based metrics and LLM-as-a-Judge scoring, featuring the novel AutoEval-Update mechanism that enables dynamic, LLM-human collaborative optimization. Our approach innovatively unifies multi-source data-driven keyword customization with adaptive prompt engineering. Extensive offline and online experiments demonstrate its effectiveness: customized ads achieve a 9% lift in click-through rate (CTR), a 12% increase in impression volume, and a 0.38% reduction in cost-per-click (CPC). AutoEval-Main attains 89.57% agreement with human evaluations, significantly improving both assessment efficiency and consistency.
๐ Abstract
Offsite marketing is essential in e-commerce, enabling businesses to reach customers through external platforms and drive traffic to retail websites. However, most current offsite marketing content is overly generic, template-based, and poorly aligned with landing pages, limiting its effectiveness. To address these limitations, we propose MarketingFM, a retrieval-augmented system that integrates multiple data sources to generate keyword-specific ad copy with minimal human intervention. We validate MarketingFM via offline human and automated evaluations and large-scale online A/B tests. In one experiment, keyword-focused ad copy outperformed templates, achieving up to 9% higher CTR, 12% more impressions, and 0.38% lower CPC, demonstrating gains in ad ranking and cost efficiency. Despite these gains, human review of generated ads remains costly. To address this, we propose AutoEval-Main, an automated evaluation system that combines rule-based metrics with LLM-as-a-Judge techniques to ensure alignment with marketing principles. In experiments with large-scale human annotations, AutoEval-Main achieved 89.57% agreement with human reviewers. Building on this, we propose AutoEval-Update, a cost-efficient LLM-human collaborative framework to dynamically refine evaluation prompts and adapt to shifting criteria with minimal human input. By selectively sampling representative ads for human review and using a critic LLM to generate alignment reports, AutoEval-Update improves evaluation consistency while reducing manual effort. Experiments show the critic LLM suggests meaningful refinements, improving LLM-human agreement. Nonetheless, human oversight remains essential for setting thresholds and validating refinements before deployment.