Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data

📅 2026-06-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of commonly reported point estimates—such as F1 scores—in text classification, which often lack reliable uncertainty quantification, particularly in settings involving small samples, rare classes, or nested data structures (e.g., texts nested within individuals). The authors systematically evaluate multiple confidence interval methods and propose a bootstrap-based F1 estimator augmented with pseudocount regularization. They further demonstrate that accurate inference in nested designs requires simultaneous adjustment of both effective sample size and degrees of freedom. Empirical results show that conventional Wald intervals suffer from undercoverage, whereas the recommended Agresti–Coull, Wilson, and hierarchical bootstrap methods substantially improve coverage accuracy, offering a more robust approach to uncertainty quantification in domains such as the social sciences.
📝 Abstract
Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language, providing metrics such as recall and precision as evidence of their validity. Yet, though these metrics are point estimates subject to sampling variation, measures of uncertainty are inconsistently reported alongside them. Further, when they are reported, they are often estimated with methods that are not appropriate when relevant labelled datasets are small or performance is high. To increase and improve confidence interval reporting in the field, this paper evaluates confidence interval methods for performance metrics under conditions typical of social science text classification: small to moderate sample sizes, infrequent constructs, and texts nested within individuals. Across simulations, default methods such as the Wald interval and the basic percentile bootstrap are the least accurate, with coverage sometimes far below the nominal 95% level. Accuracy is improved with the use of Agresti-Coull, Wilson, Clopper-Pearson, and a novel pseudo-count regularized bootstrap (which is particularly relevant to the calculation of F1). When texts are nested within individuals, we demonstrate that adjustment for both effective N and the appropriate degrees of freedom is necessary for producing accurate analytic intervals. Among bootstrap intervals, the hierarchical bootstrap is more accurate than the cluster bootstrap when individuals produce a moderate number of texts but overly conservative when individuals produce only a few. By providing guidance to the field on appropriate interval estimation, we aim to improve the transparency of machine learning applications, and to encourage greater attention to the validation sample size at the design stage.
Problem

Research questions and friction points this paper is trying to address.

uncertainty estimation
classifier performance
confidence intervals
nested data
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

confidence intervals
nested data
text classification
large language models
bootstrap methods
K
Kylie Anglin
Department of Educational Psychology, Neag School of Education, University of Connecticut