When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis

📅 2026-05-27
📈 Citations: 0
Influential: 0
📄 PDF

career value

176K/year
🤖 AI Summary
Current evaluation methods struggle to uncover substantial disagreements among large language models (LLMs) in public opinion classification, potentially misleading policy decisions. This work proposes an interpretability-focused auditing framework that treats inter-model disagreement as a signal of semantic complexity, directing human review toward genuinely ambiguous opinions. Through multi-model comparisons, expert-defined scoring rules, and a two-stage annotation experiment, the study finds that thematic disagreements across models significantly outweigh variations caused by prompt perturbations within a single model. While expert rules mitigate superficial discrepancies, they fail to resolve deeper cognitive divergences. Moreover, human annotators frequently introduce novel interpretive frameworks absent from model outputs. Moving beyond conventional accuracy metrics, this paradigm highlights the diagnostic value of disagreement in interpretive coding for nuanced opinion analysis.
📝 Abstract
Federal agencies are deploying large language models (LLMs) to categorize public comment corpora, where the model's organization of the record shapes what policymakers see and which arguments register. Standard evaluation, anchored on stance accuracy against a small validated set, cannot detect when different models produce materially different categorizations of the same public input. We propose an Interpretive Audit Pipeline that treats multi-model disagreement as diagnostic of interpretive complexity and directs human review toward genuinely ambiguous public input. Analyzing 1,260 public comments on a federal USDA docket across four LLMs, we find that inter-model thematic divergence exceeds within-model prompt variation, and that an expert rubric suppresses deep interpretive disagreement without resolving it. In a two-stage labeling study on a stratified 40-comment subsample, four LLMs and a human annotator labeled independently and then revised after seeing the others' labels. Revision behavior varied across labelers, and the human annotator's revisions frequently introduced framings absent from the ensemble's collective output. We argue disagreement-based evaluation is a necessary complement to accuracy metrics for LLM-assisted interpretive coding.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
public comment analysis
model disagreement
interpretive coding
stance accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interpretive Audit Pipeline
LLM disagreement
public comment analysis
interpretive coding
model evaluation
🔎 Similar Papers
No similar papers found.
A
Aisha Najera
AI Lab, Princeton University; Engineering and Applied Sciences, RAND Corporation
A
Alvin Moon
Engineering and Applied Sciences, RAND Corporation
V
Vedant Srinivasan
Science, Technology, and International Affairs, Georgetown University
R
Rajesh Veeraraghavan
Science, Technology, and International Affairs, Georgetown University