Score
Generating and calibrating synthetic populations and priors to match public demographic data so planning tools and scores reflect local population variants, enabling calibrated infrastructure recommendations and practitioner-ready reports.
This study addresses the challenge of generating high-fidelity synthetic populations in the absence of microdata. We propose a novel constraint programming–based generative framework that directly encodes macro-level statistical constraints—such as cross-tabulated distributions of age, education, and occupation—as well as structural relational constraints, bypassing conventional sampling-based inference. This ensures strict consistency of individual attributes and exact global statistical alignment with target distributions. Our approach innovatively integrates constraint solving with aggregate data analysis and incorporates a large language model interface to enhance semantic modeling of categorical attributes. Empirical evaluation on official census data demonstrates that the framework robustly reproduces multidimensional statistical distributions, quantifies bias propagation into downstream policy simulations, and significantly improves reproducibility and decision reliability in social behavior modeling, market analysis, and policy evaluation.
Synthetic population models for urban social simulation suffer from inadequate privacy protection, biased underrepresentation of marginalized subpopulations (e.g., low-income or immigrant groups), and insufficient geographic fidelity. Method: We propose the first Wasserstein GAN framework for multi-city spatial synthetic population generation. It integrates EU-SILC microdata with weighted sampling, spatially constrained modeling, and statistical calibration—novelly incorporating EU-SILC survey weights and external demographic constraints to explicitly mitigate systematic underrepresentation of rare subgroups. Results: Evaluated on Helsinki and Thessaloniki, our framework generates high-fidelity synthetic populations that significantly improve distributional consistency, enhance representativeness of marginalized groups, and prevent simulated discrimination. It delivers a reproducible, scalable, and fairness-aware paradigm for privacy-preserving, agent-based urban simulation.
Traditional population synthesis methods (e.g., Iterative Proportional Fitting) suffer from degraded fidelity in high-dimensional settings and fail to capture structured intra-household dependencies; existing deep learning approaches lack controllability and explicit modeling of household-level correlations. To address these limitations, we propose the Conditional Input Directed Acyclic Table Generative Adversarial Network (ciDATGAN)—the first table-GAN framework integrating conditional generation with directed acyclic graph (DAG)-structured embeddings to explicitly model heterogeneous, asymmetric dependencies among household members. Leveraging data augmentation and multi-stage optimization, ciDATGAN synthesizes a high-fidelity population dataset comprising 20 million individuals across 7.5 million households. The synthetic data strictly preserves marginal distributions aligned with census statistics, achieves 17% higher diversity than U.S. Public Use Microdata Samples (PUMS) and 13% over the Popgen baseline, and significantly improves representational capacity and fairness in urban and transportation simulation tasks.
Existing synthetic population methods struggle to generate individual-level data with precise latitude-longitude coordinates—instead producing only coarse-grained regional aggregates—due to the sparsity and extreme skewness of geographic coordinate distributions, which are difficult to model effectively. To address this, we propose a novel NF+VAE joint framework: for the first time, we employ Normalizing Flows to map raw geographic coordinates into a regularized latent space that explicitly captures spatial autocorrelation, then integrate this with a Variational Autoencoder to jointly model the distribution of spatial and non-spatial attributes. Evaluated on 121 real-world datasets, our method generates statistically faithful and privacy-preserving fine-grained household locations, significantly outperforming copula-based and uniform allocation baselines. We further introduce a multidimensional evaluation framework balancing spatial accuracy, practical utility, and privacy protection—enabling high-resolution applications such as flood response and epidemic spread modeling.
Population synthesis for target-year scenarios in transportation and urban planning faces challenges including high-dimensional data modeling, poor scalability, and the zero-cell problem. Method: This paper proposes a hybrid population synthesis method integrating Conditional Tabular Generative Adversarial Networks (CT-GAN) with Fitness-Based Sampling Combinatorial Optimization (FBS-CO). It is the first to apply CT-GAN to target-population generation within a marginal-constraint-driven hybrid modeling framework, jointly preserving univariate distribution fidelity and multivariate relational consistency. Results: Experiments show that pure CT-GAN achieves optimal univariate distribution matching; the hybrid model significantly outperforms conventional FBS-CO in satisfying multidimensional target-year marginal constraints and exhibits strong robustness to zero-frequency cells. Consequently, it enhances the statistical validity and scenario applicability of high-dimensional synthetic populations.
Existing generative population synthesis models struggle to assess whether scenario targets are compatible with the learned joint population structure, often introducing structural biases. This work proposes an ensemble-based Bayesian updating framework that employs a population-aware conditional variational autoencoder to model the underlying structural distribution and treats scenario targets as probabilistic evidence in the form of aggregate statistics. The influence of these targets on structural uncertainty is quantified through posterior weights. Crucially, the method introduces effective sample size (ESS) as a novel metric to evaluate scenario compatibility, revealing that synthetic outcomes depend not only on the magnitude of the targets but more fundamentally on their alignment with the joint population structure. This approach provides an interpretable probabilistic diagnostic tool for assessing scenario feasibility in transportation planning and enables the identification of potential structural failure modes.
This study addresses the limitations of existing empirical risk assessment frameworks, which rely on assumptions about sample data and are ill-suited for privacy risk analysis of synthetic population-scale datasets. The authors demonstrate that conventional membership inference attacks (MIAs) may fail in full-population synthesis scenarios, necessitating a reevaluation of attribute inference and individual identifiability risks. They advocate for context-sensitive privacy evaluations grounded in specific application settings. To this end, the work critically reexamines MIAs and attribute inference attacks (AIAs), proposing a revised privacy risk assessment framework tailored to population-level synthetic data. This framework exposes fundamental shortcomings in current evaluation paradigms and provides both theoretical foundations and methodological guidance for developing next-generation privacy metrics aligned with population-scale data science.
This study investigates whether large language model (LLM)-driven synthetic populations exhibit controllability—defined as the capacity to generate ordered, reproducible responses to external stimuli of known valence that align with a predefined population structure. In the fictional city of Montelago, we constructed 120 synthetic agents with explicit latent structures and evaluated their reactions to seven institutionally framed messages of varying valence using generative synthetic populations (GSP), LLM-based agents, a pre-registered validation framework (SIVE), and temperature parameter sweeps. All seven pre-registered metrics were satisfied across all temperature settings; a “weakly positive” message initially misclassified as negative was identified and corrected. Instrument noise remained stable at approximately half the magnitude of inter-agent variability, and individual-level trajectories revealed dynamic patterns obscured in aggregate statistics. The work establishes controllability as a core criterion for internal validity in synthetic populations and repurposes calibration failures into a diagnostic tool uncovering unexpected interactions between message sentiment and agent trust.
This study addresses the underestimation of variance and inferential bias in synthetic data arising from informative sampling and missing data in complex surveys. The authors propose a Bayesian synthesis framework that simultaneously imputes missing values and generates synthetic data through an adaptive weighting mechanism. By integrating principles of informative sampling theory within a Bayesian modeling paradigm, the method ensures consistent parameter estimation while yielding an asymptotically efficient Godambe information matrix. This overcomes the systematic underestimation of uncertainty inherent in conventional Bayesian synthesis approaches. Simulation studies demonstrate that the proposed method accurately quantifies uncertainty for both model parameters and population-level inferences, substantially enhancing the statistical reliability of synthetic datasets.
Traditional surveys are costly, time-consuming, and challenging to control for demographic variables, while existing simulation methods often produce inauthentic responses due to insufficient modeling of structured individual backgrounds. To address these limitations, this work proposes a large-scale virtual survey simulation platform designed for non-technical users, which uniquely integrates the Anthology and Alterity frameworks. By leveraging structured narrative contexts to guide large language models, the platform generates demographically consistent responses and supports open-ended generation, probabilistic resampling, and multimodal inputs—including text, images, and audio. Evaluated on tasks ranging from political typology and biomedical topics to preference elicitation for New Yorker cartoon captions, the platform yields opinion distributions that significantly outperform baseline methods and closely align with real-world data, demonstrating its validity and practical utility.