🤖 AI Summary
This study addresses the challenge of simultaneously preserving marginal distributions and inter-column dependencies in mixed-type tabular data generation by proposing the MIND model. Its core innovation lies in decoupling marginal modeling from dependency learning: column-wise optimal transport maps heterogeneous features into a unified latent space, while a conditional diffusion model captures cross-column dependencies. Furthermore, copula-based tangent denoising and rank projection during sampling are introduced to effectively suppress marginal shifts. Experiments across nine benchmark datasets demonstrate that MIND significantly improves both marginal fidelity and dependency preservation, achieving an excellent balance between generative quality and downstream predictive utility.
📝 Abstract
This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.