🤖 AI Summary
This study addresses the challenge of generating high-fidelity synthetic populations in the absence of microdata. We propose a novel constraint programming–based generative framework that directly encodes macro-level statistical constraints—such as cross-tabulated distributions of age, education, and occupation—as well as structural relational constraints, bypassing conventional sampling-based inference. This ensures strict consistency of individual attributes and exact global statistical alignment with target distributions. Our approach innovatively integrates constraint solving with aggregate data analysis and incorporates a large language model interface to enhance semantic modeling of categorical attributes. Empirical evaluation on official census data demonstrates that the framework robustly reproduces multidimensional statistical distributions, quantifies bias propagation into downstream policy simulations, and significantly improves reproducibility and decision reliability in social behavior modeling, market analysis, and policy evaluation.
📝 Abstract
We introduce a constraint-programming framework for generating synthetic populations that reproduce target statistics with high precision while enforcing full individual consistency. Unlike data-driven approaches that infer distributions from samples, our method directly encodes aggregated statistics and structural relations, enabling exact control of demographic profiles without requiring any microdata. We validate the approach on official demographic sources and study the impact of distributional deviations on downstream analyses. This work is conducted within the Pollitics project developed by Emotia, where synthetic populations can be queried through large language models to model societal behaviors, explore market and policy scenarios, and provide reproducible decision-grade insights without personal data.