Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks
This work addresses the challenge of annotation disagreement in subjective NLP tasks, where ambiguity in labeling criteria or overlapping category boundaries often leads to inconsistent judgments. The authors propose a pattern-level diagnostic framework that introduces an interpretable, criterion-level auditing mechanism prior to label aggregation. By collecting fine-grained evaluations from multiple annotators on individual labeling guidelines, the method systematically identifies two failure modes: unstable annotation standards and systematic category overlap. This approach provides the first structured attribution of disagreement early in the annotation pipeline, offering empirical grounding for refining annotation schemes. Evaluated on a commercial document task involving persuasion-value extraction, the analysis reveals that disagreements concentrate around a few unstable criteria, with nearly half of the sentences activating multiple categories. The diagnostic outcomes show strong alignment with domain expert assessments and effectively inform revisions to the annotation guidelines.