Reproducibility is not construct validity: LLM measurement of institutionally situated communication

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用欧盟AI法案咨询数据,通过比较调查回应和LLM注释的自由文本提交,探讨了高注释可重复性并不等同于结构有效性的问题。
📝 Abstract
High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.
Problem

Research questions and friction points this paper is trying to address.

reproducibility
construct validity
large language models
stakeholder groups
AI risks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reproducibility
Construct Validity
LLM Annotation
Stakeholder Groups
Spatial Autocorrelation
V
Veronika Batzdorfer
Department of Sociology and Computational Sociology, Karlsruhe Institute of Technology (KIT), Douglasstr. 24, 76133 Karlsruhe, Germany
C
Carlo Romano Marcello Alessandro Santagiustina
Inria Paris Centre, Inria, 48 rue Barrault, 75013 Paris, France; médialab, Sciences Po, 1 Saint-Thomas, 75007 Paris, France