Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation
This study addresses the instability of large language model (LLM) annotations under varying task designs, revealing substantial “instrument uncertainty” that confidence scores fail to mitigate. By evaluating seven LLMs across twelve task designs on a sample of 3,000 tweets, this work systematically quantifies the impact of prompt engineering on annotation outcomes using Fleiss’ and Cohen’s Kappa coefficients. The findings demonstrate that design variations inflate the variance of prevalence estimates by 76- to 110-fold, far exceeding typical inter-annotator disagreement among humans. Furthermore, this research establishes that comparing multiple task designs is the only effective approach for measuring such uncertainty, exposing critical limitations in existing evaluation methodologies and highlighting systematic biases introduced by both model selection and prompt design.