Variance-Aware LLM Annotation for Strategy Research: Sources, Diagnostics, and a Protocol for Reliable Measurement

📅 2025-12-02
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the pronounced instability of large language models (LLMs) in textual annotation for strategic research, demonstrating that neglecting their inherent variance can lead to non-reproducible findings and measurement bias. The work systematically identifies five sources of variance in LLM-based annotation and integrates content analysis, generalizability theory, prompt engineering, model selection, and statistical aggregation to develop a structured, variance-aware annotation protocol. This framework delineates the applicability boundaries of LLM annotation, offers an auditable measurement system, and optimizes aggregation rules and reporting standards under constrained sampling budgets. Empirical results reveal that minor design variations can induce result fluctuations of 12–85 percentage points, whereas the proposed approach substantially enhances annotation reliability and research reproducibility.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Analogy

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability theory, we diagnose five variance sources: construct specification, interface effects, model preferences, output extraction, and system-level aggregation. Empirical demonstrations show that minor design choices-prompt phrasing, model selection-can shift outcomes by 12-85 percentage points. Such variance threatens not only reproducibility but econometric identification: annotation errors correlated with covariates bias parameter estimates regardless of average accuracy. We develop a variance-aware protocol specifying sampling budgets, aggregation rules, and reporting standards, and delineate scope conditions where LLM annotation should not be used. These contributions transform LLM-based annotation from ad hoc practice into auditable measurement infrastructure.
Problem

Research questions and friction points this paper is trying to address.

variance
LLM annotation
reproducibility
measurement error
strategy research
Innovation

Methods, ideas, or system contributions that make the work stand out.

variance-aware annotation
LLM reliability
measurement protocol
annotation bias
generalizability theory
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arnaldo Camuffo
Bocconi University, ION Management Science Lab.
Alfonso Gambardella
Alfonso Gambardella
Bocconi University, Milan
strategyinnovationapplied economics
S
Saeid Kazemi
Bocconi University, ION Management Science Lab.
J
Jakub Malachowski
Bocconi University, ION Management Science Lab.
A
Abhinav Pandey
Bocconi University, ION Management Science Lab.