🤖 AI Summary
This study addresses the lack of identifiability theory for causal representation learning under weak intervention assumptions and the challenge of unsupervised visual state estimation. We propose a block-structured disentanglement framework grounded in realistic intervention mechanisms, establishing identifiability theory for block-wise disentanglement under weak assumptions and applying it to robotic visual state estimation. This work bridges rigorous theoretical foundations and practical applications by enabling the direct recovery of physical variables from images. Experimental validation in controlled embodied scenarios demonstrates that the proposed method achieves high-fidelity latent variable recovery without labeled data, effectively closing the gap between causal representation learning and embodied intelligence applications.
📝 Abstract
Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which temper the direct application to real-world problems. This paper has a two-fold objective for interventional CRL. First, it establishes identifiability guarantees for substantially weaker interventional assumptions, resulting in block disentanglement of the causal variables, where the block structure depends on the realistically available intervention mechanisms. Secondly, the block disentanglement framework is used for embodied visual state estimation, in which the objective is to recover the latent physical variables of a robotic system directly from visual data (images and videos) without labeled data. These two components are critically complementary. The block disentanglement theory delineates identifiability guarantees under weakened assumptions, and the application demonstrates that the resulting objective remains effective in a controlled embodied setting despite further assumption violations, providing a theory-to-practice bridge needed to translate the promise of label-free CRL into practical problems.