🤖 AI Summary
Current evaluation methods struggle to uncover substantial disagreements among large language models (LLMs) in public opinion classification, potentially misleading policy decisions. This work proposes an interpretability-focused auditing framework that treats inter-model disagreement as a signal of semantic complexity, directing human review toward genuinely ambiguous opinions. Through multi-model comparisons, expert-defined scoring rules, and a two-stage annotation experiment, the study finds that thematic disagreements across models significantly outweigh variations caused by prompt perturbations within a single model. While expert rules mitigate superficial discrepancies, they fail to resolve deeper cognitive divergences. Moreover, human annotators frequently introduce novel interpretive frameworks absent from model outputs. Moving beyond conventional accuracy metrics, this paradigm highlights the diagnostic value of disagreement in interpretive coding for nuanced opinion analysis.
📝 Abstract
Federal agencies are deploying large language models (LLMs) to categorize public comment corpora, where the model's organization of the record shapes what policymakers see and which arguments register. Standard evaluation, anchored on stance accuracy against a small validated set, cannot detect when different models produce materially different categorizations of the same public input. We propose an Interpretive Audit Pipeline that treats multi-model disagreement as diagnostic of interpretive complexity and directs human review toward genuinely ambiguous public input. Analyzing 1,260 public comments on a federal USDA docket across four LLMs, we find that inter-model thematic divergence exceeds within-model prompt variation, and that an expert rubric suppresses deep interpretive disagreement without resolving it. In a two-stage labeling study on a stratified 40-comment subsample, four LLMs and a human annotator labeled independently and then revised after seeing the others' labels. Revision behavior varied across labelers, and the human annotator's revisions frequently introduced framings absent from the ensemble's collective output. We argue disagreement-based evaluation is a necessary complement to accuracy metrics for LLM-assisted interpretive coding.