🤖 AI Summary
This study addresses the pronounced instability of large language models (LLMs) in textual annotation for strategic research, demonstrating that neglecting their inherent variance can lead to non-reproducible findings and measurement bias. The work systematically identifies five sources of variance in LLM-based annotation and integrates content analysis, generalizability theory, prompt engineering, model selection, and statistical aggregation to develop a structured, variance-aware annotation protocol. This framework delineates the applicability boundaries of LLM annotation, offers an auditable measurement system, and optimizes aggregation rules and reporting standards under constrained sampling budgets. Empirical results reveal that minor design variations can induce result fluctuations of 12–85 percentage points, whereas the proposed approach substantially enhances annotation reliability and research reproducibility.
📝 Abstract
Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability theory, we diagnose five variance sources: construct specification, interface effects, model preferences, output extraction, and system-level aggregation. Empirical demonstrations show that minor design choices-prompt phrasing, model selection-can shift outcomes by 12-85 percentage points. Such variance threatens not only reproducibility but econometric identification: annotation errors correlated with covariates bias parameter estimates regardless of average accuracy. We develop a variance-aware protocol specifying sampling budgets, aggregation rules, and reporting standards, and delineate scope conditions where LLM annotation should not be used. These contributions transform LLM-based annotation from ad hoc practice into auditable measurement infrastructure.