Score
Designs and implements methods that discover, represent, and evaluate human‑interpretable concepts or categories from unlabeled data using techniques such as clustering, matrix factorization (e.g., NMF), prototype or part discovery, and latent‑space induction. Builds tools to extract concept spaces and concept activations, align discovered concepts to semantic labels, detect novel categories, and analyze interpretability and concept quality (coherence, completeness, separability).
Social science research urgently requires interpretable and reproducible exploratory discovery from unstructured text—without presupposing measurement constructs. This paper proposes an end-to-end framework addressing this challenge. First, it constructs a high-dimensional, semantically transparent, and interpretable concept dictionary via sparse coding and semantic modeling. Second, it introduces a novel high-dimensional multiple testing procedure that rigorously controls the k-familywise error rate (k-FWER) under arbitrary variable dependence, substantially reducing researcher degrees of freedom. Third, it integrates machine learning interpretability techniques with selective inference to ensure statistical validity. The method is empirically validated in economic analyses—both causal and descriptive—and is accompanied by an open-source Jupyter toolkit, enabling low-cost, fully reproducible empirical workflows.
This work addresses the limitations of existing concept-based models, which rely on fine-grained annotations and treat concepts as flat, independent units, thereby hindering the construction of interpretable, hierarchical concept structures. To overcome this, the authors propose Multi-Level Concept Segmentation (MLCS) and Deep Hierarchical Concept Embedding Models (Deep-HiCEMs), which require only coarse-grained top-level supervision to automatically discover multi-layered, human-interpretable concept hierarchies. The framework supports concept interventions across abstraction levels and successfully uncovers novel, explainable concepts absent from training data across multiple benchmarks. While maintaining high predictive accuracy, the method significantly enhances task performance through test-time interventions, marking the first approach capable of automatically constructing a multi-granular concept system from coarse-grained labels alone.
Weak interpretability of large language models (LLMs), unreliability of post-hoc explanation methods, and limitations of concept bottleneck models (CBMs)—including dependence on costly human annotations, restricted representational capacity, and lack of interpretability at the task level—motivate this work. We propose the self-supervised Interpretable Concept Embedding Model (ICEM), the first framework to introduce concept modeling into the textual domain. ICEM leverages the inherent generalization capability of LLMs to autonomously predict concept labels without manual annotation. It enables end-to-end interpretable prediction via concept embeddings and an interpretable decision function, supporting concept intervention, logical attribution, and controllable decoding-path steering. On text classification tasks, ICEM achieves performance comparable to fully supervised CBMs and black-box LLMs, while providing human-understandable, causally grounded explanations. Thus, ICEM unifies interpretability, interactivity, and controllability in a single architecture.
This work addresses the lack of human-interpretable concepts in intermediate-layer representations of CNNs. We propose an unsupervised post-hoc method that optimizes an orthogonal rotation in feature space to extract disentangled, concept-level interpretable basis vectors from sparsely thresholded activation responses. Unlike supervised approaches relying on manually annotated concepts, ours is the first purely unsupervised paradigm for discovering highly interpretable bases. We further introduce an improved interpretability metric and a concept-alignment analysis framework, validating our method across multiple CNN architectures and datasets. Experiments demonstrate that the rotated intermediate representations significantly outperform supervised basis extraction methods in both conceptual diversity and interpretability. Our results reveal an inherent limitation of supervised paradigms—namely, their restricted coverage of conceptual breadth—and open a new direction for model interpretability research. (149 words)
High-stakes domains—such as healthcare and finance—demand interpretable clustering outcomes to ensure transparency, accountability, and regulatory compliance. Method: This survey systematically analyzes over 120 scholarly works, proposing the first unified taxonomy of interpretability dimensions for clustering. It rigorously distinguishes intrinsically interpretable models—including rule-based, prototype-based, and sparsity-driven approaches—from post-hoc explanation techniques—such as visualization, feature attribution, and local surrogate modeling. The study further develops a use-case-oriented, structured classification framework and principled evaluation criteria. Contribution/Results: It introduces the first practical guideline for selecting appropriate interpretable clustering methods based on application requirements. The work bridges theoretical foundations with real-world deployment, providing both conceptual clarity and actionable insights to support the development and adoption of clustering algorithms that jointly optimize accuracy and interpretability—thereby advancing trustworthy AI in ethically and regulatorily sensitive contexts.
研究提出了一种探索性非结构化数据分析框架,通过用户研究和CLIP评估,识别了四个提高人机协作的机会。
This work addresses the lack of semantic interpretability in existing novel class discovery methods by embedding the task into a structured semantic concept space. Leveraging pretrained multimodal models, it aligns visual–language similarity priors to learn concept representations from unlabeled data and jointly optimizes self-labeling objectives for both labeled and unlabeled samples in the concept-space logits. The approach inherently endows each discovered class with human-interpretable explanations composed of stable concept signatures and instance-level evidence. Theoretically, it constrains the hypothesis space and favors semantically coherent partitions. Experiments demonstrate state-of-the-art performance on CIFAR-10 (92.63%), CIFAR-100 (76.45%), and CUB-200, with the added distinction of being the only method that provides both cluster-level and instance-level human-readable interpretations in novel class discovery.
This work addresses the challenge of ambiguous fine-grained category boundaries in generalized category discovery (GCD), which arises from reliance solely on visual information and the decoupling of supervision from the discovery process. To this end, we propose the Analogical Text Concept Generator (ATCG), which introduces an analogical reasoning mechanism to generate textual concepts for unlabeled samples based on known categories. By fusing these textual concepts with visual features, ATCG reformulates category discovery as a joint vision–language reasoning process, effectively transferring prior knowledge to enhance discriminability. Designed as a plug-and-play module, ATCG is compatible with both parametric and clustering-based GCD methods without requiring modifications to existing pipelines. Extensive experiments across six benchmark datasets demonstrate consistent and significant improvements in overall, known-class, and novel-class performance, with the largest gains observed in fine-grained scenarios.
Existing concept embedding methods struggle to model inter-concept relationships and rely heavily on multi-granularity human annotations, limiting both interpretability and practical applicability. This work proposes Hierarchical Concept Embedding Models (HiCEMs), which introduce hierarchical structure into concept representations for the first time and integrate an unsupervised concept splitting technique to automatically discover fine-grained subconcepts from pretrained models. HiCEMs generate multi-level, interpretable embeddings without requiring additional annotations and support test-time multi-granularity interventions. Evaluated across multiple datasets—including the newly introduced PseudoKitchens—the approach demonstrates its ability to uncover human-understandable subconcepts, improve task accuracy, and enable effective explanation and intervention.
本文探讨使用程序作为概念的通用表示方法,以解决人类如何从稀疏数据中学习和泛化新概念的问题。