π€ AI Summary
This study addresses the limitation that training causal normalizing flows typically requires complete data, rendering them difficult to apply directly to datasets with missing values. To overcome this, we propose MissCNF, a model trained by maximizing the marginal likelihood of observed data without imputation or sample discarding. Leveraging an autoregressive structure, it integrates out missing variables exclusively within their ancestral closure. Furthermore, we introduce a "causal family positivity" condition to ensure distributional identifiability under extreme missingness, accommodating MCAR, MAR, and MNAR mechanisms. Empirically, MissCNF significantly outperforms conventional methods in KL divergence and counterfactual error under nonlinear settings while maintaining state-of-the-art performance in linear scenarios, demonstrating robustness even with up to 90% missing data.
π Abstract
Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only missing variables in the ancestral closure of the observed set are integrated out, while the others are dropped without computation. We further establish the conditions under which MissCNF recovers the true joint distribution, and introduce \emph{causal-family positivity}, where identification is possible even when no record in the dataset is ever complete. We compare MissCNF with two common strategies for handling missing data: listwise deletion and impute-then-fit pipelines. Across eight synthetic causal benchmarks, three missingness mechanisms, and missing rates up to $90\%$, MissCNF achieves the lowest KL divergence in 23 of 24 nonlinear MCAR and MAR settings and in all nonlinear MNAR settings, as well as the lowest counterfactual error in 20 of 24 settings. On linear SCMs, where linear imputation performs best, MissCNF ranks in the top two in 22 of 24 settings.