Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks

📅 2026-04-18
📈 Citations: 0
Influential: 0
📄 PDF

career value

132K/year
🤖 AI Summary
This work addresses the challenge of annotation disagreement in subjective NLP tasks, where ambiguity in labeling criteria or overlapping category boundaries often leads to inconsistent judgments. The authors propose a pattern-level diagnostic framework that introduces an interpretable, criterion-level auditing mechanism prior to label aggregation. By collecting fine-grained evaluations from multiple annotators on individual labeling guidelines, the method systematically identifies two failure modes: unstable annotation standards and systematic category overlap. This approach provides the first structured attribution of disagreement early in the annotation pipeline, offering empirical grounding for refining annotation schemes. Evaluated on a commercial document task involving persuasion-value extraction, the analysis reveals that disagreements concentrate around a few unstable criteria, with nearly half of the sentences activating multiple categories. The diagnostic outcomes show strong alignment with domain expert assessments and effectively inform revisions to the annotation guidelines.

Technology Category

Application Category

📝 Abstract
Subjective NLP datasets typically aggregate annotator judgments into a single gold label, making it difficult to diagnose whether disagreement reflects unclear criteria, collapsed distinctions, or legitimate plurality. We propose a \emph{schema-level diagnostic} for auditing expert-designed annotation schemas \emph{prior to} gold-label commitment, using only multi-annotator criterion judgments. The diagnostic separates two failure modes: unstable criteria with hard-to-operationalize boundaries, and systematic overlap that blurs the boundaries between mutually exclusive categories. Applied to persuasive value extraction in commercial documents, we find that disagreement is not diffuse: instability concentrates in a few criteria, while nearly half of covered sentences activate multiple categories. These signals align with where domain experts disagree, yielding an evidence-based audit for tightening guidelines, revising category structure, or reconsidering the annotation paradigm.
Problem

Research questions and friction points this paper is trying to address.

subjective NLP
annotation disagreement
gold label
annotation schema
criterion instability
Innovation

Methods, ideas, or system contributions that make the work stand out.

schema-level diagnostic
annotation disagreement
subjective NLP
criterion instability
category overlap
🔎 Similar Papers
No similar papers found.