🤖 AI Summary
This paper investigates intervention-driven causal representation learning (CRL) for nonparametric latent causal models, where the mapping from latent variables to observations may be linear or nonlinear. It addresses identifiability and algorithmic realizability of both latent causal variables and the underlying causal graph. Methodologically, it establishes the first theoretical connection between score functions (i.e., gradients of log-densities) and CRL; introduces general identifiability conditions that do not require intervention environment labels—significantly relaxing standard intervention assumptions; and proposes a score-matching-based optimization framework supporting both stochastic hard and soft interventions, as well as latent graph structure recovery. Theoretically, full identifiability is guaranteed with a single hard intervention in the linear case and two hard interventions in the general nonlinear setting. Extensive experiments on synthetic and image data demonstrate accurate recovery of both latent causal variables and the causal graph structure.
📝 Abstract
This paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transformation that maps the latent variables to the observed variables. Linear and general transformations are investigated. The paper addresses both the identifiability and achievability aspects. Identifiability refers to determining algorithm-agnostic conditions that ensure recovering the true latent causal variables and the latent causal graph underlying them. Achievability refers to the algorithmic aspects and addresses designing algorithms that achieve identifiability guarantees. By drawing novel connections between score functions (i.e., the gradients of the logarithm of density functions) and CRL, this paper designs a score-based class of algorithms that ensures both identifiability and achievability. First, the paper focuses on linear transformations and shows that one stochastic hard intervention per node suffices to guarantee identifiability. It also provides partial identifiability guarantees for soft interventions, including identifiability up to ancestors for general causal models and perfect latent graph recovery for sufficiently non-linear causal models. Secondly, it focuses on general transformations and shows that two stochastic hard interventions per node suffice for identifiability. Notably, one does not need to know which pair of interventional environments have the same node intervened. Finally, the theoretical results are empirically validated via experiments on structured synthetic data and image data.