DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited interpretability and generalizability of existing concept models that rely on fixed categories or linguistic supervision by proposing DisParQ, a self-supervised discrete concept learning framework built upon a frozen DINOv2 backbone. Pioneering a label-free and language-free paradigm, DisParQ achieves unsupervised, spatially grounded part-level concept and quantized attribute learning through prototype dictionary assignment, sparse activation, continuous residual quantization, and spatial decoding reconstruction. Experimental results demonstrate that DisParQ attains 83.2% accuracy on ImageNet linear probing while surpassing language-aligned models in concept consistency. Furthermore, it enables fine-grained cross-category part retrieval, effectively enhancing the interpretability of visual representations.
📝 Abstract
Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone. It requires no class labels and no language supervision. Each image patch is assigned to exactly one concept from a learnable prototype dictionary, and only a sparse subset of concepts may activate per image. To capture how each concept varies across images (e.g., the type of a "wheel"), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes. A spatial decoder reconstructs the backbone's representation from the concepts and attributes alone, so successful reconstruction means that the discrete representation preserves the backbone's information. We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers). We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.
Problem

Research questions and friction points this paper is trying to address.

interpretable vision models
concept-based representation
self-supervised learning
discrete concepts
vision foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised learning
Discrete concepts
Vision foundation models
Quantized attributes
Interpretability
🔎 Similar Papers
No similar papers found.