When AI Agents Disagree Like Humans: Reasoning Trace Analysis for Human-AI Collaborative Moderation

📅 2026-04-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work challenges the conventional view in multi-agent systems that treats disagreement among AI agents as mere noise to be eliminated, arguing instead that such divergence may reflect genuine value pluralism—particularly in culturally and subjectively charged tasks like hate speech moderation. The study proposes a novel, reasoning-structure-based taxonomy classifying AI disagreements into four types and implements it using five large language model (LLM) agents with diverse perspectives, generating reasoning traces on the Measuring Hate Speech dataset. By combining embedding and classification models, the framework identifies patterns such as “convergent disagreement.” Experimental results show that when agent conclusions align, human annotator disagreement drops significantly (d > 0.8), and the structure of AI disagreements strongly correlates with human conflicts. These findings demonstrate the approach’s efficacy in guiding human–AI collaborative decision-making and advocate for shifting multi-agent systems from consensus-seeking toward explicit uncertainty representation.

Technology Category

Multiagent Systems: Agreement, Argumentation & NegotiationHumans and AI: Learning Human Values and PreferencesCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsResponsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
When LLM-based multi-agent systems disagree, current practice treats this as noise to be resolved through consensus. We propose it can be signal. We focus on hate speech moderation, a domain where judgments depend on cultural context and individual value weightings, producing high legitimate disagreement among human annotators. We hypothesize that convergent disagreement, where agents reason similarly but conclude differently, indicates genuine value pluralism that humans also struggle to resolve. Using the Measuring Hate Speech corpus, we embed reasoning traces from five perspective-differentiated agents and classify disagreement patterns using a four-category taxonomy based on reasoning similarity and conclusion agreement. We find that raw reasoning divergence weakly predicts human annotator conflict, but the structure of agent discord carries additional signal: cases where agents agree on a verdict show markedly lower human disagreement than cases where they do not, with large effect sizes (d>0.8) surviving correction for multiple comparisons. Our taxonomy-based ordering correlates with human disagreement patterns. These preliminary findings motivate a shift from consensus-seeking to uncertainty-surfacing multi-agent design, where disagreement structure - not magnitude - guides when human judgment is needed.
Problem

Research questions and friction points this paper is trying to address.

multi-agent disagreement
hate speech moderation
value pluralism
reasoning trace
human-AI collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

reasoning trace analysis
multi-agent disagreement
value pluralism
human-AI collaboration
hate speech moderation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Michał Wawer
Faculty of Electronics and Information Technology, Warsaw University of Technology, Warsaw, Poland
J
Jarosław A. Chudziak
Faculty of Electronics and Information Technology, Warsaw University of Technology, Warsaw, Poland