unsupervised concept explainability

Designs and implements methods that discover, represent, and evaluate human‑interpretable concepts or categories from unlabeled data using techniques such as clustering, matrix factorization (e.g., NMF), prototype or part discovery, and latent‑space induction. Builds tools to extract concept spaces and concept activations, align discovered concepts to semantic labels, detect novel categories, and analyze interpretability and concept quality (coherence, completeness, separability).

unsupervisedconceptexplainability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Social science research urgently requires interpretable and reproducible exploratory discovery from unstructured text—without presupposing measurement constructs. This paper proposes an end-to-end framework addressing this challenge. First, it constructs a high-dimensional, semantically transparent, and interpretable concept dictionary via sparse coding and semantic modeling. Second, it introduces a novel high-dimensional multiple testing procedure that rigorously controls the k-familywise error rate (k-FWER) under arbitrary variable dependence, substantially reducing researcher degrees of freedom. Third, it integrates machine learning interpretability techniques with selective inference to ensure statistical validity. The method is empirically validated in economic analyses—both causal and descriptive—and is accompanied by an open-source Jupyter toolkit, enabling low-cost, fully reproducible empirical workflows.

Developing statistically principled discovery framework for unstructured dataEnabling replicable unsupervised analysis with minimal researcher degreesPerforming interpretable high-dimensional hypothesis testing on concepts

This work addresses the limitations of existing concept-based models, which rely on fine-grained annotations and treat concepts as flat, independent units, thereby hindering the construction of interpretable, hierarchical concept structures. To overcome this, the authors propose Multi-Level Concept Segmentation (MLCS) and Deep Hierarchical Concept Embedding Models (Deep-HiCEMs), which require only coarse-grained top-level supervision to automatically discover multi-layered, human-interpretable concept hierarchies. The framework supports concept interventions across abstraction levels and successfully uncovers novel, explainable concepts absent from training data across multiple benchmarks. While maintaining high predictive accuracy, the method significantly enhances task performance through test-time interventions, marking the first approach capable of automatically constructing a multi-granular concept system from coarse-grained labels alone.

concept hierarchyconcept-based modelshierarchical representation

Self-supervised Interpretable Concept-based Models for Text Classification

Jun 20, 2024
FD
Francesco De Santis
🏛️ Politecnico di Torino | Università della Svizzera Italiana | University of Cambridge

Weak interpretability of large language models (LLMs), unreliability of post-hoc explanation methods, and limitations of concept bottleneck models (CBMs)—including dependence on costly human annotations, restricted representational capacity, and lack of interpretability at the task level—motivate this work. We propose the self-supervised Interpretable Concept Embedding Model (ICEM), the first framework to introduce concept modeling into the textual domain. ICEM leverages the inherent generalization capability of LLMs to autonomously predict concept labels without manual annotation. It enables end-to-end interpretable prediction via concept embeddings and an interpretable decision function, supporting concept intervention, logical attribution, and controllable decoding-path steering. On text classification tasks, ICEM achieves performance comparable to fully supervised CBMs and black-box LLMs, while providing human-understandable, causally grounded explanations. Thus, ICEM unifies interpretability, interactivity, and controllability in a single architecture.

Enhancing interpretability of text analysis modelsOvercoming limitations of Concept-Bottleneck Models (CBMs)Reducing dependency on extensive concept annotations

Unsupervised Interpretable Basis Extraction for Concept-Based Visual Explanations

Mar 19, 2023
AD
Alexandros Doumanoglou
🏛️ Information Technologies Institute (ITI) | Centre for Research and Technology HELLAS (CERTH) | University of Maastricht

This work addresses the lack of human-interpretable concepts in intermediate-layer representations of CNNs. We propose an unsupervised post-hoc method that optimizes an orthogonal rotation in feature space to extract disentangled, concept-level interpretable basis vectors from sparsely thresholded activation responses. Unlike supervised approaches relying on manually annotated concepts, ours is the first purely unsupervised paradigm for discovering highly interpretable bases. We further introduce an improved interpretability metric and a concept-alignment analysis framework, validating our method across multiple CNN architectures and datasets. Experiments demonstrate that the rotated intermediate representations significantly outperform supervised basis extraction methods in both conceptual diversity and interpretability. Our results reveal an inherent limitation of supervised paradigms—namely, their restricted coverage of conceptual breadth—and open a new direction for model interpretability research. (149 words)

Enhancing interpretability of intermediate layer representations through unsupervised basis transformationExtracting interpretable basis directions from CNN feature spaces without concept supervisionIdentifying interpretable feature directions collectively using sparsity optimization instead of annotations

Interpretable Clustering: A Survey

Sep 01, 2024
LH
Lianyu Hu
🏛️ Dalian University of Technology

High-stakes domains—such as healthcare and finance—demand interpretable clustering outcomes to ensure transparency, accountability, and regulatory compliance. Method: This survey systematically analyzes over 120 scholarly works, proposing the first unified taxonomy of interpretability dimensions for clustering. It rigorously distinguishes intrinsically interpretable models—including rule-based, prototype-based, and sparsity-driven approaches—from post-hoc explanation techniques—such as visualization, feature attribution, and local surrogate modeling. The study further develops a use-case-oriented, structured classification framework and principled evaluation criteria. Contribution/Results: It introduces the first practical guideline for selecting appropriate interpretable clustering methods based on application requirements. The work bridges theoretical foundations with real-world deployment, providing both conceptual clarity and actionable insights to support the development and adoption of clustering algorithms that jointly optimize accuracy and interpretability—thereby advancing trustworthy AI in ethically and regulatorily sensitive contexts.

Addressing the trade-off between clustering accuracy and interpretabilityDeveloping a taxonomy to classify explainable clustering algorithmsProviding transparent clustering methods for high-stakes application domains

Latest Papers

What's happening recently
View more

This work addresses the lack of semantic interpretability in existing novel class discovery methods by embedding the task into a structured semantic concept space. Leveraging pretrained multimodal models, it aligns visual–language similarity priors to learn concept representations from unlabeled data and jointly optimizes self-labeling objectives for both labeled and unlabeled samples in the concept-space logits. The approach inherently endows each discovered class with human-interpretable explanations composed of stable concept signatures and instance-level evidence. Theoretically, it constrains the hypothesis space and favors semantically coherent partitions. Experiments demonstrate state-of-the-art performance on CIFAR-10 (92.63%), CIFAR-100 (76.45%), and CUB-200, with the added distinction of being the only method that provides both cluster-level and instance-level human-readable interpretations in novel class discovery.

Explainable AIInterpretabilityNovel Category Discovery

This work addresses the challenge of ambiguous fine-grained category boundaries in generalized category discovery (GCD), which arises from reliance solely on visual information and the decoupling of supervision from the discovery process. To this end, we propose the Analogical Text Concept Generator (ATCG), which introduces an analogical reasoning mechanism to generate textual concepts for unlabeled samples based on known categories. By fusing these textual concepts with visual features, ATCG reformulates category discovery as a joint vision–language reasoning process, effectively transferring prior knowledge to enhance discriminability. Designed as a plug-and-play module, ATCG is compatible with both parametric and clustering-based GCD methods without requiring modifications to existing pipelines. Extensive experiments across six benchmark datasets demonstrate consistent and significant improvements in overall, known-class, and novel-class performance, with the largest gains observed in fine-grained scenarios.

analogical reasoningfine-grained categoriesGeneralized Category Discovery

Existing concept embedding methods struggle to model inter-concept relationships and rely heavily on multi-granularity human annotations, limiting both interpretability and practical applicability. This work proposes Hierarchical Concept Embedding Models (HiCEMs), which introduce hierarchical structure into concept representations for the first time and integrate an unsupervised concept splitting technique to automatically discover fine-grained subconcepts from pretrained models. HiCEMs generate multi-level, interpretable embeddings without requiring additional annotations and support test-time multi-granularity interventions. Evaluated across multiple datasets—including the newly introduced PseudoKitchens—the approach demonstrates its ability to uncover human-understandable subconcepts, improve task accuracy, and enable effective explanation and intervention.

annotation burdenConcept Embedding Modelsconcept relationships

Hot Scholars

NK

Nishanth Kumar

Ph.D. Student, MIT
Artificial IntelligenceMachine LearningRoboticsPlanning
TS

Tom Silver

Assistant Professor at Princeton
PlanningLearningRobotics
KK

Kristian Kersting

Professor of AI & ML, Technical University of Darmstadt, Hessian.ai, DFKI, CAIRNE/ELLIS, AAAI Fellow
Artificial IntelligenceNeurosymbolic AIProbabilistic CircuitsMachine Learning
JF

Jonas Fischer

Group Leader, Max-Planck-Institute for Informatics
Machine LearningXAIComputational Biology
SS

Sanchit Sinha

University of Virginia
Natural Language ProcessingMachine LearningComputer Vision