Score
Design and build models, annotation procedures, and evaluation methods that identify when one speaker or text expresses disagreement with another, classifying utterances, reply pairs, or document-level interactions as disagreeing versus not; and analyze linguistic cues, conversational structure, pragmatics, and model errors to improve detection accuracy and robustness.
NLP models often suffer from reduced reliability and trustworthiness in practice due to reliance on or generation of conflicting information—arising from factual contradictions and subjective biases in natural text, annotator disagreements and societal biases in training data, and hallucinations or knowledge inconsistencies during model interaction. This paper provides the first unified taxonomy of conflicts across the entire NLP pipeline, introduces a cross-scenario conflict classification framework and an extensible mitigation paradigm, and fills a critical gap in systematic surveys on conflict-aware modeling. By integrating semantic consistency analysis, annotation robustness evaluation, generation credibility calibration, and multi-perspective reasoning, we establish a human-in-the-loop mechanism for conflict detection and resolution. The work clarifies the fundamental impact of conflicts on model trustworthiness, identifies six core challenges, and outlines future research directions—thereby offering both theoretical foundations and practical guidelines for developing interpretable, debuggable, and trustworthy NLP systems.
Current evaluation methods struggle to uncover substantial disagreements among large language models (LLMs) in public opinion classification, potentially misleading policy decisions. This work proposes an interpretability-focused auditing framework that treats inter-model disagreement as a signal of semantic complexity, directing human review toward genuinely ambiguous opinions. Through multi-model comparisons, expert-defined scoring rules, and a two-stage annotation experiment, the study finds that thematic disagreements across models significantly outweigh variations caused by prompt perturbations within a single model. While expert rules mitigate superficial discrepancies, they fail to resolve deeper cognitive divergences. Moreover, human annotators frequently introduce novel interpretive frameworks absent from model outputs. Moving beyond conventional accuracy metrics, this paradigm highlights the diagnostic value of disagreement in interpretive coding for nuanced opinion analysis.
This work investigates whether large language models (LLMs) can effectively model human annotation disagreement—a critical signal of task subjectivity and instance ambiguity. Current evaluation paradigms predominantly assess accuracy against majority-voted labels, neglecting models’ capacity to capture annotation uncertainty. To address this gap, we propose the first systematic evaluation framework for disagreement prediction grounded in single-annotator labels, integrated with RLVR-style reasoning to quantify LLMs’ fidelity to empirical annotation distributions. Our experiments reveal three key findings: (1) mainstream LLMs exhibit poor calibration in predicting human disagreement; (2) majority-label accuracy substantially obscures this limitation; and (3) incorporating reinforcement learning–based reasoning degrades disagreement prediction performance, exposing a misalignment between standard optimization objectives and uncertainty modeling. We publicly release our code and datasets to advance more holistic, human-centered LLM evaluation.
This study investigates inappropriate directed language—encompassing explicit hate speech and implicit discriminatory expressions—targeting individuals or groups in English Reddit conversations. We propose a multi-source annotation framework that systematically compares expert, crowdsourced, and ChatGPT annotators for the first time. Our context-sensitive, fine-grained taxonomy introduces novel target categories such as “social beliefs” and “body image,” and quantifies the critical influence of linguistic context on annotation decisions. Results demonstrate that ChatGPT significantly underperforms humans in detecting subtle discriminatory language; cross-source consistency analysis further identifies key sources of annotation bias. The work contributes an interpretable, inclusive annotation paradigm for automated content moderation, alongside empirically grounded pathways for model improvement. (138 words)
This work addresses the challenge of modeling and evaluating AI systems’ capacity to capture human judgment variability—such as disagreement and subjectivity. Methodologically, we (1) extend the LeWiDi benchmark to four tasks (paraphrase identification, irony/sarcasm detection, natural language inference) with ordinal annotations and individual-perspective prediction; (2) introduce the first integration of soft-label learning and annotator modeling, moving beyond hard-classification paradigms; and (3) propose a multi-task training framework jointly optimizing distributional prediction, individual annotator modeling, and population-level judgment distribution learning. Contributions include two novel evaluation metrics that surpass conventional measures like cross-entropy, and comprehensive empirical analysis revealing strengths and limitations of existing approaches in modeling judgment variability. These advances significantly enhance LeWiDi’s utility and extensibility as a benchmark platform for controversy-aware AI.
This work addresses the substantial and sample-dependent disagreement among human annotators when labeling inappropriate language, such as offensive or hateful content—a phenomenon that is difficult to quantify. The authors propose an “opposition index” to characterize the degree of annotator polarization and develop methods to predict this index and the associated annotation variance based on textual features. They systematically compare two approaches: direct regression to predict variance and variance estimation derived from predicted probability distributions. Both achieve moderate predictive performance. The study further reveals that samples with high opposition indices are more challenging for models to classify accurately and tend to have their toxicity systematically underestimated. This research offers a novel perspective and practical tools for understanding and modeling subjective annotation disagreement in toxic language detection.
This work addresses the challenge of annotator disagreement in subjective and ambiguous natural language processing tasks—such as toxicity detection and stance analysis—where divergent perspectives are often dismissed as noise rather than meaningful signals. The authors propose a domain-agnostic taxonomy of annotation disagreement alongside a unified modeling framework that explicitly captures structural relationships among annotators and supports multi-target prediction. By introducing disagreement-aware evaluation metrics, the study advocates a paradigm shift from consensus-based learning toward perspectivist modeling, offering a normative lens for fairness assessment. The paper systematically integrates existing disagreement-aware methodologies, clarifies the trajectory of this evolving paradigm, and outlines promising future directions, including the incorporation of multi-source variability and the development of interpretable disagreement frameworks.
This study addresses the challenging and highly subjective task of annotating evaluative language by focusing on the Attitude subsystem within Appraisal Theory, specifically examining expressions of affect, judgment, and appreciation in TED Talks. Through a systematic comparison of annotation behaviors among trained linguists, novice linguists, and large language models (LLMs), it reveals for the first time the performance disparities between human expertise and LLMs in such nuanced tasks. By integrating three prompt engineering strategies with model fine-tuning, the research significantly enhances the LLM’s capacity for automatic classification of Appraisal categories, achieving an F1 score of 0.77 post-fine-tuning—surpassing novice linguists and approaching the performance of expert linguists. These findings demonstrate the promising potential of LLMs as effective assistants in complex theoretical annotation tasks within digital humanities.
This work addresses the lack of systematic evaluation of large language models’ responses to user disagreement, particularly their susceptibility to undesirable conversational tendencies such as flattery or stubbornness. The authors propose a feedback metric framework based on fictitious response-rebuttal (FR) pairings, enabling the first quantitative assessment of model flattery and stubbornness in multiple-choice settings without ground-truth answers. This approach facilitates cross-model and cross-topic comparisons of dialogic behavior. Empirical evaluation on physics-related questions reveals that newer OpenAI models and higher reasoning configurations significantly reduce flattery, thereby demonstrating the framework’s effectiveness and generalizability.
This study addresses the challenges of misclassification and limited local practitioner buy-in in online polarization and hate speech monitoring, stemming from cultural contextual differences. To tackle these issues, the project collaborates with peacebuilders and data scientists from Kenya and Sudan through a participatory annotation process to jointly define tasks, design labeling schemes, and iteratively validate models. The methodology deeply integrates domain experts throughout the AI development lifecycle, combining BERT-based fine-tuned classifiers with a cross-cultural contextual validation mechanism. The resulting open-source classifiers—Kenya-polarization and Sudan-hate-speech—released on Hugging Face, demonstrate significantly improved classification accuracy under culturally sensitive conditions and greater local acceptance of the tools, achieving synergistic optimization of technical robustness, contextual validity, and ethical alignment.