🤖 AI Summary
This study addresses the limitations of conventional synthetic data generation, including reliance on iterative training, difficulties in cross-modal adaptation, and opaque generation processes, by proposing GENSCRIPT, a training-free inference framework. The method leverages large language models for semantic reasoning and employs coding agents to compile statistical profiles and constraints into auditable sampler code, thereby unifying the synthesis of tabular, time-series, and relational data. Its core contribution lies in pioneering a training-free paradigm that ensures logical consistency and behavioral interpretability. Remarkably, the framework requires only two minutes to construct and six seconds to sample 50,000 rows, while faithfully preserving primary-foreign key relationships and marginal distributions. Empirical evaluations demonstrate that GENSCRIPT significantly outperforms existing baselines.
📝 Abstract
Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.