demographic calibration

Generating and calibrating synthetic populations and priors to match public demographic data so planning tools and scores reflect local population variants, enabling calibrated infrastructure recommendations and practitioner-ready reports.

demographiccalibration

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of generating high-fidelity synthetic populations in the absence of microdata. We propose a novel constraint programming–based generative framework that directly encodes macro-level statistical constraints—such as cross-tabulated distributions of age, education, and occupation—as well as structural relational constraints, bypassing conventional sampling-based inference. This ensures strict consistency of individual attributes and exact global statistical alignment with target distributions. Our approach innovatively integrates constraint solving with aggregate data analysis and incorporates a large language model interface to enhance semantic modeling of categorical attributes. Empirical evaluation on official census data demonstrates that the framework robustly reproduces multidimensional statistical distributions, quantifies bias propagation into downstream policy simulations, and significantly improves reproducibility and decision reliability in social behavior modeling, market analysis, and policy evaluation.

Enables societal modeling without requiring personal microdataGenerates synthetic populations with exact demographic controlStudies impact of distributional deviations on downstream analyses

Synthetic population models for urban social simulation suffer from inadequate privacy protection, biased underrepresentation of marginalized subpopulations (e.g., low-income or immigrant groups), and insufficient geographic fidelity. Method: We propose the first Wasserstein GAN framework for multi-city spatial synthetic population generation. It integrates EU-SILC microdata with weighted sampling, spatially constrained modeling, and statistical calibration—novelly incorporating EU-SILC survey weights and external demographic constraints to explicitly mitigate systematic underrepresentation of rare subgroups. Results: Evaluated on Helsinki and Thessaloniki, our framework generates high-fidelity synthetic populations that significantly improve distributional consistency, enhance representativeness of marginalized groups, and prevent simulated discrimination. It delivers a reproducible, scalable, and fairness-aware paradigm for privacy-preserving, agent-based urban simulation.

Privacy ProtectionUrban Population ModelingWasserstein Generative Adversarial Networks

Deep and diverse population synthesis for multi-person households using generative models

Aug 13, 2025
HY
Hai Yang
🏛️ New York University Tandon School of Engineering

Traditional population synthesis methods (e.g., Iterative Proportional Fitting) suffer from degraded fidelity in high-dimensional settings and fail to capture structured intra-household dependencies; existing deep learning approaches lack controllability and explicit modeling of household-level correlations. To address these limitations, we propose the Conditional Input Directed Acyclic Table Generative Adversarial Network (ciDATGAN)—the first table-GAN framework integrating conditional generation with directed acyclic graph (DAG)-structured embeddings to explicitly model heterogeneous, asymmetric dependencies among household members. Leveraging data augmentation and multi-stage optimization, ciDATGAN synthesizes a high-fidelity population dataset comprising 20 million individuals across 7.5 million households. The synthetic data strictly preserves marginal distributions aligned with census statistics, achieves 17% higher diversity than U.S. Public Use Microdata Samples (PUMS) and 13% over the Popgen baseline, and significantly improves representational capacity and fairness in urban and transportation simulation tasks.

Capturing associations among household members accuratelyGenerating high-dimensional synthetic household population dataImproving diversity and equity in population synthesis

Population synthesis with geographic coordinates

Oct 08, 2025
JL
Jacopo Lenti
🏛️ Sapienza University of Rome | CENTAI Institute | BIFI Institute | University of Zaragoza | Intesa Sanpaolo Innovation Center | Intesa Sanpaolo

Existing synthetic population methods struggle to generate individual-level data with precise latitude-longitude coordinates—instead producing only coarse-grained regional aggregates—due to the sparsity and extreme skewness of geographic coordinate distributions, which are difficult to model effectively. To address this, we propose a novel NF+VAE joint framework: for the first time, we employ Normalizing Flows to map raw geographic coordinates into a regularized latent space that explicitly captures spatial autocorrelation, then integrate this with a Variational Autoencoder to jointly model the distribution of spatial and non-spatial attributes. Evaluated on 121 real-world datasets, our method generates statistically faithful and privacy-preserving fine-grained household locations, significantly outperforming copula-based and uniform allocation baselines. We further introduce a multidimensional evaluation framework balancing spatial accuracy, practical utility, and privacy protection—enabling high-resolution applications such as flood response and epidemic spread modeling.

Addressing uneven spatial density and empty areas in coordinatesGenerating synthetic populations with precise geographic coordinatesLearning joint distribution of spatial and non-spatial features

Target Population Synthesis using CT-GAN

Oct 01, 2025
TR
Tanay Rastogi

Population synthesis for target-year scenarios in transportation and urban planning faces challenges including high-dimensional data modeling, poor scalability, and the zero-cell problem. Method: This paper proposes a hybrid population synthesis method integrating Conditional Tabular Generative Adversarial Networks (CT-GAN) with Fitness-Based Sampling Combinatorial Optimization (FBS-CO). It is the first to apply CT-GAN to target-population generation within a marginal-constraint-driven hybrid modeling framework, jointly preserving univariate distribution fidelity and multivariate relational consistency. Results: Experiments show that pure CT-GAN achieves optimal univariate distribution matching; the hybrid model significantly outperforms conventional FBS-CO in satisfying multidimensional target-year marginal constraints and exhibits strong robustness to zero-frequency cells. Consequently, it enhances the statistical validity and scenario applicability of high-dimensional synthetic populations.

Generating target populations for transportation scenario planningIntegrating deep generative models with traditional optimization methodsOvercoming deterministic synthesis limitations with high-dimensional data

Latest Papers

What's happening recently
View more

Existing generative population synthesis models struggle to assess whether scenario targets are compatible with the learned joint population structure, often introducing structural biases. This work proposes an ensemble-based Bayesian updating framework that employs a population-aware conditional variational autoencoder to model the underlying structural distribution and treats scenario targets as probabilistic evidence in the form of aggregate statistics. The influence of these targets on structural uncertainty is quantified through posterior weights. Crucially, the method introduces effective sample size (ESS) as a novel metric to evaluate scenario compatibility, revealing that synthetic outcomes depend not only on the magnitude of the targets but more fundamentally on their alignment with the joint population structure. This approach provides an interpretable probabilistic diagnostic tool for assessing scenario feasibility in transportation planning and enables the identification of potential structural failure modes.

Bayesian evaluationgenerative population synthesisscenario compatibility

This study addresses the limitations of existing empirical risk assessment frameworks, which rely on assumptions about sample data and are ill-suited for privacy risk analysis of synthetic population-scale datasets. The authors demonstrate that conventional membership inference attacks (MIAs) may fail in full-population synthesis scenarios, necessitating a reevaluation of attribute inference and individual identifiability risks. They advocate for context-sensitive privacy evaluations grounded in specific application settings. To this end, the work critically reexamines MIAs and attribute inference attacks (AIAs), proposing a revised privacy risk assessment framework tailored to population-level synthetic data. This framework exposes fundamental shortcomings in current evaluation paradigms and provides both theoretical foundations and methodological guidance for developing next-generation privacy metrics aligned with population-scale data science.

empirical riskpopulation-level dataprivacy

This study investigates whether large language model (LLM)-driven synthetic populations exhibit controllability—defined as the capacity to generate ordered, reproducible responses to external stimuli of known valence that align with a predefined population structure. In the fictional city of Montelago, we constructed 120 synthetic agents with explicit latent structures and evaluated their reactions to seven institutionally framed messages of varying valence using generative synthetic populations (GSP), LLM-based agents, a pre-registered validation framework (SIVE), and temperature parameter sweeps. All seven pre-registered metrics were satisfied across all temperature settings; a “weakly positive” message initially misclassified as negative was identified and corrected. Instrument noise remained stable at approximately half the magnitude of inter-agent variability, and individual-level trajectories revealed dynamic patterns obscured in aggregate statistics. The work establishes controllability as a core criterion for internal validity in synthetic populations and repurposes calibration failures into a diagnostic tool uncovering unexpected interactions between message sentiment and agent trust.

controllabilityinstrument calibrationinternal validity

This study addresses the underestimation of variance and inferential bias in synthetic data arising from informative sampling and missing data in complex surveys. The authors propose a Bayesian synthesis framework that simultaneously imputes missing values and generates synthetic data through an adaptive weighting mechanism. By integrating principles of informative sampling theory within a Bayesian modeling paradigm, the method ensures consistent parameter estimation while yielding an asymptotically efficient Godambe information matrix. This overcomes the systematic underestimation of uncertainty inherent in conventional Bayesian synthesis approaches. Simulation studies demonstrate that the proposed method accurately quantifies uncertainty for both model parameters and population-level inferences, substantially enhancing the statistical reliability of synthetic datasets.

Bayesian FrameworkInformative SamplingMissing Data

Traditional surveys are costly, time-consuming, and challenging to control for demographic variables, while existing simulation methods often produce inauthentic responses due to insufficient modeling of structured individual backgrounds. To address these limitations, this work proposes a large-scale virtual survey simulation platform designed for non-technical users, which uniquely integrates the Anthology and Alterity frameworks. By leveraging structured narrative contexts to guide large language models, the platform generates demographically consistent responses and supports open-ended generation, probabilistic resampling, and multimodal inputs—including text, images, and audio. Evaluated on tasks ranging from political typology and biomedical topics to preference elicitation for New Yorker cartoon captions, the platform yields opinion distributions that significantly outperform baseline methods and closely align with real-world data, demonstrating its validity and practical utility.

backstory-conditioneddemographic controllarge language models

Hot Scholars

HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
FR

Francisco Rowe

Professor of Population Data Science, Geographic Data Science Lab
human mobilityinternal migrationregional sciencegeographic data science
JW

Jon Wakefield

Professor Statistics Biostatistics University of Washington
statisticsbiostatisticsepidemiology
DA

David A. Swanson

Population Research Center, Portland State University; Center for Studies in Demography & Ecology
Demography