๐ค AI Summary
This study addresses the limitations of existing tabular generative models, which require per-dataset training, exhibit restricted knowledge transfer, and incur high storage costs. To overcome these challenges, this work proposes a cross-dataset mixed-type diffusion model. The method pioneers defining the diffusion process directly within a mixed feature space, dynamically adapting to heterogeneous categorical domains via schema-constrained reverse parameterization. Furthermore, it employs a shared schema-aware Transformer denoiser combined with cascaded diffusion techniques to enable end-to-end joint training. Experimental results demonstrate that the proposed model achieves state-of-the-art average generation quality across seven real-world datasets using fewer parameters. Notably, following large-scale pretraining, the model exhibits significant generalization capabilities on previously unseen datasets, highlighting its potential for universal tabular data synthesis.
๐ Abstract
Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.