AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals

📅 2025-05-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This paper identifies a systematic “moderation bias” in LLM-as-a-Judge for content moderation evaluation: AI judges significantly overrate refusal responses that invoke ethical justifications—by 23.6% on average (p < 0.001)—compared to human annotators, while exhibiting no such preference for technically grounded refusals. Leveraging Chatbot Arena data, the study employs GPT-4o and Llama 3 70B as judge models, integrated with human annotations and a controlled comparative analysis framework, to first formally define and empirically validate this bias. Key contributions are threefold: (1) introducing “moderation bias” as a novel concept that exposes a value misalignment between automated evaluation metrics and human moral reasoning; (2) demonstrating that current alignment benchmarks may misguide safety-oriented training by rewarding ethically verbose but potentially superficial refusals; and (3) advocating for a paradigm shift in evaluation design—incorporating human cognitive nuance and contextual sensitivity to ensure robust, value-aligned moderation assessment.

Technology Category

Application Category

📝 Abstract
As large language models (LLMs) are increasingly deployed in high-stakes settings, their ability to refuse ethically sensitive prompts-such as those involving hate speech or illegal activities-has become central to content moderation and responsible AI practices. While refusal responses can be viewed as evidence of ethical alignment and safety-conscious behavior, recent research suggests that users may perceive them negatively. At the same time, automated assessments of model outputs are playing a growing role in both evaluation and training. In particular, LLM-as-a-Judge frameworks-in which one model is used to evaluate the output of another-are now widely adopted to guide benchmarking and fine-tuning. This paper examines whether such model-based evaluators assess refusal responses differently than human users. Drawing on data from Chatbot Arena and judgments from two AI judges (GPT-4o and Llama 3 70B), we compare how different types of refusals are rated. We distinguish ethical refusals, which explicitly cite safety or normative concerns (e.g.,"I can't help with that because it may be harmful"), and technical refusals, which reflect system limitations (e.g.,"I can't answer because I lack real-time data"). We find that LLM-as-a-Judge systems evaluate ethical refusals significantly more favorably than human users, a divergence not observed for technical refusals. We refer to this divergence as a moderation bias-a systematic tendency for model-based evaluators to reward refusal behaviors more than human users do. This raises broader questions about transparency, value alignment, and the normative assumptions embedded in automated evaluation systems.
Problem

Research questions and friction points this paper is trying to address.

Examining AI vs human evaluation of ethical refusal responses
Assessing moderation bias in LLM-as-a-Judge frameworks
Investigating transparency issues in automated content moderation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge for automated content moderation
Ethical vs. technical refusal response classification
Moderation bias in model-based evaluators
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.