🤖 AI Summary
Unsupervised interpretable learning for high-dimensional natural data (e.g., images) remains challenging due to the lack of identifiable, semantically meaningful representations.
Method: This paper models semantic concepts as discrete implicit causal variables and constructs an identifiable multilevel causal hierarchy. It formally defines discrete concepts as hierarchical causal latent variables and establishes novel identifiability conditions for continuous high-dimensional observations—enabling complex causal structures beyond trees and DAGs. The approach integrates causal representation learning, hierarchical latent modeling, identifiability analysis, and latent diffusion mechanisms.
Contributions/Results: We theoretically prove identifiability of intricate hierarchical concepts under unsupervised learning. Synthetic experiments validate both effectiveness and robustness. Furthermore, we uncover and empirically substantiate a hierarchical generative mechanism for implicit concepts within latent diffusion models—revealing their intrinsic causal organization.
📝 Abstract
Learning concepts from natural high-dimensional data (e.g., images) holds potential in building human-aligned and interpretable machine learning models. Despite its encouraging prospect, formalization and theoretical insights into this crucial task are still lacking. In this work, we formalize concepts as discrete latent causal variables that are related via a hierarchical causal model that encodes different abstraction levels of concepts embedded in high-dimensional data (e.g., a dog breed and its eye shapes in natural images). We formulate conditions to facilitate the identification of the proposed causal model, which reveals when learning such concepts from unsupervised data is possible. Our conditions permit complex causal hierarchical structures beyond latent trees and multi-level directed acyclic graphs in prior work and can handle high-dimensional, continuous observed variables, which is well-suited for unstructured data modalities such as images. We substantiate our theoretical claims with synthetic data experiments. Further, we discuss our theory's implications for understanding the underlying mechanisms of latent diffusion models and provide corresponding empirical evidence for our theoretical insights.