🤖 AI Summary
This study addresses the discrepancy wherein LLM-as-a-judge evaluations achieve ranking consistency in occupational assessments yet exhibit severe deviations from human data in acceptance rates and aggregate metrics. To investigate this, we propose an auditing framework grounded in O*NET-BENCH that systematically evaluates 33 model configurations against real worker ratings, incorporating cross-validated calibration and prediction-assisted estimation techniques. Our findings reveal a counterintuitive phenomenon: listwise prompting protocols optimize rank ordering while degrading mean alignment. Furthermore, we demonstrate that calibrated scores possess limited explanatory power, indicating that ranking consistency alone is insufficient for supporting valid occupational measurement. Consequently, this work underscores the necessity of independently validating acceptance rates and aggregate values against ground-truth distributions to ensure LLM-based assessments remain reliable in applied settings.
📝 Abstract
LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.