Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing methods in simulating coherent respondent preferences across questions within complete surveys, as they can only reproduce single-question marginal distributions. To overcome this, we propose the FR-LLM framework, which fine-tunes large language models combined with autoregressive modeling to generate full survey responses. The core innovation lies in the Marginal-Constrained Joint Projection (MCJP) technique, which pioneers projecting joint distributions into a space satisfying individual marginal constraints, thereby unifying global consistency with local precision. Experimental results demonstrate that this approach more accurately reproduces multi-question dependency patterns across diverse datasets while exhibiting strong generalizability. Furthermore, it achieves the highest profitability in business decision-making simulations.
📝 Abstract
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Survey Simulation
Respondent Coherence
Multi-question Response Patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Survey Simulation
Autoregressive Model
Marginal-Constrained Joint Projection
Cross-item Coherence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Ji Huang
School of Computer Science, University of Science and Technology of China
M
Mengfei Li
School of Management, Fudan University
Shuai Shao
Shuai Shao
University of Science and Technology of China
Complexity TheoryInformation Theory