🤖 AI Summary
This study addresses the susceptibility of large language models (LLMs) to misinterpret randomly assigned country labels as meaningful signals in social inference tasks, leading to systematic prediction biases. Employing an innovative within-record randomized audit design, the authors conduct controlled experiments across seven social indicators in six countries using five prominent API-based LLMs, generating 14,400 paired predictions compiled into the PROV-FORECAST dataset. The experimental framework integrates demographic anchors and human ground-truth responses to rigorously assess how metadata disclosure influences reliance on spurious cues. Results demonstrate that authentic country metadata significantly reduces Brier loss (−0.040); however, even when explicitly informed of their randomness, spurious country labels induce substantial directional bias, and such disclosures fail to reliably mitigate this effect—revealing a critical deficit in LLMs’ ability to robustly discern the provenance and validity of contextual metadata.
📝 Abstract
Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.