LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of unified evaluation standards and unknown user preferences regarding advertisement insertion in large language models (LLMs). To this end, we construct the first human preference benchmark that isolates ad placement as a variable. Methodologically, we employ pairwise comparison experiments and multidimensional human annotation to collect over 18,000 judgments, which are subsequently used to fine-tune Qwen3. Our findings reveal the unreliability of frontier LLMs acting as judges, demonstrate that ad disclosure significantly alters user preferences, and quantify the trade-offs between user and advertiser perspectives. Notably, the fine-tuned smaller model comprehensively outperforms zero-shot frontier LLMs on prediction tasks, establishing a reliable evaluation paradigm for ad insertion in LLMs.
📝 Abstract
Inserting advertisements (ads) into consumer-facing LLM output is emerging as a new business model, but there is little shared evidence on how such ad insertion should be evaluated or how it affects user preferences. We introduce LLMAdBench, a human-preference benchmark for studying advertising in LLM-generated content. The benchmark isolates a simple but practically important decision: given a user conversation, an LLM response, and a matched advertisement, where should the ad be placed? Our dataset compares pairs of responses that differ only in ad position while holding all other conditions fixed including the user query, base answer, advertisement, and disclosure condition. Human annotators evaluate each pair based on six criteria from both advertiser's and user's perspectives. The resulting benchmark contains more than 18000 human judgments across two disclosure conditions: explicitly labeling the ad as sponsored and merging it into the response without disclosure. We use LLMAdBench to evaluate eight frontier LLMs as preference judges and find that they are not reliable substitutes for human evaluation. Even the most stable models reverse roughly one quarter of their decisions when the presentation order is swapped, agreement across models is low, and their placement preferences differ systematically from those of human annotators. Moreover, LLMAdBench contains substantial learnable signal. In particular, a Qwen3-8B model fine-tuned on the human preferences improves substantially over its base model and outperforms all zero-shot frontier judges on the held-out prediction task. Beyond model evaluation, LLMAdBench provides quantitative evidence on the advertiser-user trade-off and shows that the sponsorship disclosure systematically changes users'preference over ad placement.
Problem

Research questions and friction points this paper is trying to address.

LLM advertising
human preference
ad placement
benchmark evaluation
sponsorship disclosure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human Preference Benchmark
LLM Advertising
Preference Alignment
LLM-as-a-Judge
Sponsorship Disclosure
💼 Related Jobs
No related jobs found.
Rui Ai
Rui Ai
Massachusetts Institute of Technology
reinforcement learninggame theory
Y
Yuqing Liu
University of Michigan, Ann Arbor
S
Sitao Qiu
Shanghai Jiao Tong University
Y
Yun Qiao
The Ohio State University
Y
Yuhan Wang
University of Hong Kong
J
Jessica Xiwen Wang
McGill University
Y
Yiqi Yang
University of Chicago
L
Lihong Huang
Tongji University
R
Ruiyao Sun
Beijing Normal University
Kaifeng Zhang
Kaifeng Zhang
Columbia University
RoboticsPhysics SimulationMachine LearningComputer Vision
S
Shengze Ding
Shanghai Jiao Tong University
J
Jiaqi He
Shanghai Jiao Tong University
X
Xinman Wang
Shanghai Jiao Tong University
T
Tianhao Gao
Shanghai Jiao Tong University
J
Jimmy Qin
University of Texas at Dallas
Jianghao Lin
Jianghao Lin
Shanghai Jiao Tong University
Large Language ModelsAI AgentsRecommender Systems
C
Chonghuan Wang
University of Texas at Dallas