LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena

📅 2025-01-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how ethical content refusal by large language models (LLMs) on moral and sensitive topics affects real-user satisfaction. Using nearly 50,000 response pairs from Chatbot Arena, we construct a fine-grained refusal attribution framework to quantify user acceptance loss due to ethical refusal—finding refused responses are selected only 26% as often as standard ones. We fine-tune RoBERTa for automated refusal reason classification and integrate human annotation with large-scale statistical modeling to identify moderators: topic sensitivity, response length, and context. Counterintuitively, refusals on highly sensitive illegal topics exhibit higher win rates, while LLM-as-a-judge evaluation significantly attenuates the penalty for refusal—revealing systematic evaluator bias. Our core contributions are: (1) the first user-preference-oriented empirical paradigm for analyzing ethical refusal; (2) identification of the “refusal does not necessarily degrade preference” phenomenon; and (3) data-driven insights for jointly optimizing safety and usability.

Technology Category

Application Category

📝 Abstract
LLM safety and ethical alignment are widely discussed, but the impact of content moderation on user satisfaction remains underexplored. To address this, we analyze nearly 50,000 Chatbot Arena response-pairs using a novel fine-tuned RoBERTa model, that we trained on hand-labeled data to disentangle refusals due to ethical concerns from other refusals due to technical disabilities or lack of information. Our findings reveal a significant refusal penalty on content moderation, with users choosing ethical-based refusals roughly one-fourth as often as their preferred LLM response compared to standard responses. However, the context and phrasing play critical roles: refusals on highly sensitive prompts, such as illegal content, achieve higher win rates than less sensitive ethical concerns, and longer responses closely aligned with the prompt perform better. These results emphasize the need for nuanced moderation strategies that balance ethical safeguards with user satisfaction. Moreover, we find that the refusal penalty is notably lower in evaluations using the LLM-as-a-Judge method, highlighting discrepancies between user and automated assessments.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
User Satisfaction
Ethical Handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Content Moderation Impact
User Preference
Ethical Review
🔎 Similar Papers
No similar papers found.