LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the entanglement of objects and attributes in unsupervised image representation learning and the suboptimality of existing uniform decomposition strategies. To this end, it proposes a probabilistic generative model grounded in the Linear Representation Hypothesis (LRH). Methodologically, this work pioneers the introduction of LRH into slot space and designs a block attention architecture. Through variational inference, it jointly optimizes the evidence lower bound (ELBO) across object and attribute spaces, thereby achieving their decoupled learning. Experimental results demonstrate that the proposed method surpasses state-of-the-art approaches in terms of the Disentanglement, Completeness, and Informativeness (DCI) metric across multiple benchmark datasets. These findings effectively validate the plausibility of applying LRH within slot space and further support semantically interpretable image editing.
📝 Abstract
This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

disentangled representation
unsupervised attribute discovery
slot-based representation
Linear Representation Hypothesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Representation Hypothesis
Slot-based representation
Unsupervised disentanglement
Block attention
Probabilistic model
💼 Related Jobs
No related jobs found.
S
Sanket Gandhi
Yardi School of AI, IIT Delhi
U
Utkarsh Giri
Yardi School of AI, IIT Delhi
V
Varun Subramanium
Department of Computer Science and Engineering, IIT Delhi
Rohan Paul
Rohan Paul
Yardi School of AI, IIT Delhi; Department of Computer Science and Engineering, IIT Delhi
Parag Singla
Parag Singla
Indian Institute of Technology Delhi
Neuro-Symbolic ReasoningMachine LearningArtificial Intelligence