🤖 AI Summary
This study evaluates whether large language models (LLMs) can automatically generate questionnaires capable of effectively measuring social attitudes and rivaling established expert-designed scales. Through a within-subjects experimental design, the performance of GPT-4–generated questionnaires—elicited via structured prompts—was systematically compared against validated human-crafted scales across three domains: climate change, immigration, and diversity and inclusion. This work presents the first multi-domain, within-participant comparison between LLM-generated instruments and standard psychometric scales. Results indicate that LLM-generated questionnaires reliably capture major attitudinal divides and are suitable for exploratory, large-scale attitude assessment. However, they exhibit lower resolution in uncovering belief structures and reduced precision in differentiating subpopulations compared to expert-developed scales, suggesting promising yet supplementary utility in social science research.
📝 Abstract
Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate parts of this process by generating surveys at scale, raising questions about the comparability of such instruments to literature-grounded, human-designed surveys. We present a controlled empirical comparison between GPT-generated surveys and established survey baselines across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). GPT-generated surveys were produced using a fixed prompting framework enforcing a 3x3 structure over beliefs, perceptions, and behaviors, while human baselines were assembled from validated instruments to match survey length and construct coverage. We collected responses from U.S.-based participants, who completed both survey types, allowing direct within-subject comparison. We analyze differences in response distributions, clustering behavior, and alignment with self-identified stances. Our results show that GPT-generated surveys capture the same dominant attitudinal divisions as human-designed instruments, while exhibiting differences in the resolution of belief structure and group separation. These findings suggest that LLM-generated surveys are suited for exploratory and large-scale analyses, and can be used to complement expert-designed instruments.