Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems

📅 2026-04-22
📈 Citations: 0
Influential: 0
📄 PDF

career value

211K/year
🤖 AI Summary
This study addresses the limitation of prevailing generative AI evaluations that treat models as isolated predictors, thereby neglecting their sociotechnical nature as dynamically co-constructed within diverse cultural contexts. The work proposes reconceptualizing generative AI as a machine–society–human (MaSH) recursively co-constituted system, shifting evaluative focus from static outputs to value-laden interactions. Methodologically, evaluation is reframed as an embodied, recursive process through a distributive benchmark grounded in the World Values Survey, incorporating structured prompt sets and anchor-aware scoring, alongside a participatory realist methodology. Empirical analysis reveals value drift in early GPT-3 versions and demonstrates, in a real estate scenario, that the proposed framework effectively captures the dynamic societal impacts of generative AI, surpassing the constraints of conventional static benchmarks.

Technology Category

Application Category

📝 Abstract
In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: benchmarks do not just measure, they shape what models appear to be. Functionalist benchmarks treat models as isolated predictors, while prescriptive approaches assess what systems ought to be. Both obscure the sociotechnical processes through which meaning and values are enacted, risking the reification of narrow cultural perspectives in pluralist contexts. This thesis advances a descriptive alternative. It argues that generative AI must be evaluated as a pluralist sociotechnical system and develops Machine-Society-Human (MaSH) Loops, a framework for tracing how models, users, and institutions recursively co-construct meaning and values. Evaluation shifts from judging outputs to examining how values are enacted in interaction. Three contributions follow. Conceptually, MaSH Loops reframes evaluation as recursive, enactive process. Methodologically, the World Values Benchmark introduces a distributional approach grounded in World Values Survey data, structured prompt sets, and anchor-aware scoring. Empirically, the thesis demonstrates these through two cases: value drift in early GPT-3 and sociotechnical evaluation in real estate. A final chapter draws on participatory realism to argue that prompting and evaluation are constitutive interventions, not neutral observations. The thesis argues that static benchmarks are insufficient for generative AI. Responsible evaluation requires pluralist, process-oriented frameworks that make visible whose values are enacted. Evaluation is therefore a site of governance, shaping how AI systems are understood, deployed, and trusted.
Problem

Research questions and friction points this paper is trying to address.

generative AI
evaluation
sociotechnical systems
pluralism
values
Innovation

Methods, ideas, or system contributions that make the work stand out.

sociotechnical evaluation
MaSH Loops
pluralist AI
World Values Benchmark
value enactment
🔎 Similar Papers
No similar papers found.
R
Rebecca L. Johnson
The School of History and Philosophy of Science, Faculty of Science, The University of Sydney