🤖 AI Summary
This study addresses the fairness deficiencies of large language models (LLMs) in ad relevance judgment. It proposes a counterfactual evaluation framework leveraging GPT-4o and Qwen-7B, combining real-world log sampling with synthetic query experiments to systematically quantify multidimensional biases arising from advertiser identity, linguistic attributes, and demographic features, while exploring mitigation strategies during both inference and training phases. The findings reveal context-specific stereotyping patterns and demonstrate that input features significantly distort relevance judgments. Furthermore, this work establishes that the effectiveness of mitigation interventions is highly contingent upon information relevance and underlying data distributions, offering critical insights for developing more equitable LLM-based advertising systems.
📝 Abstract
Large language models (LLMs) are increasingly used to judge how well an advertisement matches a query, but the fairness of these judgments has received limited attention. We conduct a systematic study of fairness in relevance judgments made by LLMs for queries and advertisements. Our counterfactual framework examines the effects of advertiser identity and possible popularity, input language, and demographic wording. We study GPT-4o as a categorical relevance judge and a Qwen-7B model trained specifically for relevance prediction. The advertiser and language experiments use query and advertisement pairs sampled from real advertising logs. Controlled synthetic queries are used to study demographic associations in employment, housing, and credit. For both models, changing the advertiser identity or input language can alter the relevance assessment. Selected demographic comparisons also show patterns consistent with common stereotypes, particularly those involving gender and occupation. We further study mitigation during model inference and training. The results indicate that its effectiveness depends on whether advertiser information is relevant to the query and how advertiser labels are distributed in the training data. These findings can help advertising practitioners identify fairness risks and develop suitable mitigation methods for LLM relevance systems.