Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data

📅 2026-06-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional social science experiments rely on human subjects, which are costly, inefficient, and susceptible to sampling bias, thereby hindering accurate estimation of statistical quantities such as conditional expectations. This work formalizes pre-trained large language models (LLMs) as misspecified function estimators and establishes, for the first time, their risk equivalence to the Bayes optimal estimator within a restricted function class. By introducing a decomposition framework that separates representation bias from optimization error—and leveraging Le Cam’s two-point method, Pinsker’s inequality, and finite-sample concentration bounds—the study rigorously demonstrates that, under appropriate range conditions and proper model calibration, LLMs can asymptotically approach the Bayes optimal risk for estimating the conditional mean of human responses, substantially reducing reliance on expensive human experimentation.
📝 Abstract
Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent estimators of conditional expectations under squared loss, establishing restricted functional risk equivalence: under squared loss, the LLM induces an estimator whose risk matches the Bayes optimal risk for squared-loss prediction of conditional expectations for any inference that depends on the data only through the conditional mean. We formalize the LLM as a misspecified functional estimator $T(\hat{P}_n)$ trained on i.i.d.\ data, decompose the estimation error into representation bias $ε_{\mathrm{rep}}$ and optimization error, and prove that under mild regularity conditions the LLM's expected error converges to the irreducible population variance plus the squared representation bias, with the representation bias bounded by the Pinsker inequality. The identifiability error $δ$ propagates into the effective bias, inflating the asymptotic risk floor. We establish restricted functional risk equivalence via a bidirectional Le Cam deficiency analysis: the forward deficiency vanishes asymptotically while the reverse deficiency is exactly zero. We provide finite-sample concentration bounds and a calibration protocol with explicit decision rules. The result is a precise, provable statement: a well-calibrated LLM achieves the Bayes-optimal risk for conditional-mean-dependent inference, bounded by explicit scope conditions. In practical applications, this means that under satisfied conditions and well-calibrated models, large language models can be used in many prediction and decision-making tasks that originally relied on human experiments, approximating near-optimal statistical inference at lower cost.
Problem

Research questions and friction points this paper is trying to address.

human-response data
sampling bias
conditional expectation
statistical estimation
Bayes optimal risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
statistical estimation
risk equivalence
conditional expectation
Bayes optimality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haobo Yang
Department of Computer Science and Engineering, SUSTech University