unsupervised feature mining

Designs and implements unsupervised methods to discover and characterize interpretable features in neural network activations by mining activation geometry (directions, clusters, manifolds) and extracting feature vectors or readout directions without labeled supervision. Builds analyses and procedures to approximate features as activation directions, quantify how those features change or influence downstream readouts, and produce human‑interpretable reasoning features derived from activations.

unsupervisedfeaturemining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing interpretability methods that rely on human-annotated concepts, which are prone to bias and struggle to uncover reasoning features in large language models in an unsupervised manner. The authors propose MAG, an unsupervised framework that prepends a unified natural language instruction to inputs and measures its impact on the model’s activation space to automatically extract reasoning-relevant semantic features—without requiring labeled data. They find that certain features can be approximated by single activation directions, enabling intervention via vector arithmetic. Through activation geometry analysis, activation differencing, vector manipulation, and RFD similarity metrics, the extracted features accurately predict the model’s world knowledge. In prompt injection classification tasks, selecting training data using RFD yields 94.7% Top-1 and 100% Top-2 accuracy.

activation geometryinterpretabilitylarge language models

Unsupervised Interpretable Basis Extraction for Concept-Based Visual Explanations

Mar 19, 2023
AD
Alexandros Doumanoglou
🏛️ Information Technologies Institute (ITI) | Centre for Research and Technology HELLAS (CERTH) | University of Maastricht

This work addresses the lack of human-interpretable concepts in intermediate-layer representations of CNNs. We propose an unsupervised post-hoc method that optimizes an orthogonal rotation in feature space to extract disentangled, concept-level interpretable basis vectors from sparsely thresholded activation responses. Unlike supervised approaches relying on manually annotated concepts, ours is the first purely unsupervised paradigm for discovering highly interpretable bases. We further introduce an improved interpretability metric and a concept-alignment analysis framework, validating our method across multiple CNN architectures and datasets. Experiments demonstrate that the rotated intermediate representations significantly outperform supervised basis extraction methods in both conceptual diversity and interpretability. Our results reveal an inherent limitation of supervised paradigms—namely, their restricted coverage of conceptual breadth—and open a new direction for model interpretability research. (149 words)

Enhancing interpretability of intermediate layer representations through unsupervised basis transformationExtracting interpretable basis directions from CNN feature spaces without concept supervisionIdentifying interpretable feature directions collectively using sparsity optimization instead of annotations

Knowledge Discovery using Unsupervised Cognition

Sep 30, 2024
AI
Alfredo Ibias
🏛️ Avatar Cognition

To address challenges in unsupervised knowledge discovery—including weak modeling of feature correlations, poor pattern interpretability, and semantic distortion in dimensionality reduction—this paper proposes a three-stage analytical framework grounded in an Unsupervised Cognition model: (1) association pattern mining, (2) cognition-significance-driven interpretable feature selection, and (3) semantic-consistency-constrained dimensionality reduction. It pioneers the end-to-end integration of cognitive modeling with knowledge discovery, enabling fully label-free, interpretable analysis. Evaluated on diverse multi-source empirical datasets, the method consistently outperforms state-of-the-art approaches across all core metrics: pattern completeness (+12.7%), feature discriminability (+9.4%), and dimensionality-reduction interpretability (+15.3%).

Data Complexity ReductionHidden PatternsUnsupervised Learning

Refining Neural Activation Patterns for Layer-Level Concept Discovery in Neural Network-Based Receivers

May 21, 2025
MT
Marko Tuononen
🏛️ Nokia Networks | Nokia Bell Labs | University of Eastern Finland

This paper addresses the challenge of identifying hierarchical, distributed activation patterns in neural networks. To overcome limitations of neuron-level or hand-crafted interpretable feature analyses, we propose Neural Activation Pattern (NAP) modeling based on full-layer activation distributions. Methodologically, we introduce normalized preprocessing, kernel density estimation for distribution modeling, SNR-adaptive distance metrics, and hierarchical clustering. For the first time, our approach uncovers a continuous activation manifold in neural communication receivers that is dominantly governed by signal-to-noise ratio (SNR), demonstrating its physical interpretability. Experiments show that NAP significantly improves separation between in-distribution and out-of-distribution samples, empirically confirming SNR as a critical implicit factor learned by the model. The method enhances generalization, interpretability, and reliability diagnostics—enabling robust, physics-informed analysis of deep neural representations.

Analyzing SNR's role in shaping activation manifolds in receiver modelsIdentifying layer-level concepts in neural networks via activation patternsImproving NAP methodology for better concept discovery and generalization

Decision Trees for Interpretable Clusters in Mixture Models and Deep Representations

Nov 03, 2024
MF
Maximilian Fleissner
🏛️ Technical University of Munich

This work addresses a fundamental challenge in interpretable clustering: whether the worst-case interpretability cost bound can be surpassed—and the underlying cluster structure reliably recovered—when data exhibit well-separated clusters. To this end, we propose a decision-tree-based mixture-model clustering method. We introduce the first theoretical framework of “interpretability–noise ratio,” enabling data-agnostic, efficient tree construction. Under sub-Gaussian assumptions, we derive tight upper and lower bounds on estimation error. Furthermore, we pioneer the integration of Concept Activation Vectors (CAVs) into unsupervised clustering, facilitating interpretable cluster identification in deep representation spaces. Experiments on standard tabular and image benchmarks demonstrate that our method significantly enhances interpretability while maintaining high clustering accuracy; moreover, its theoretical guarantees strictly improve upon those of existing distribution-agnostic approaches.

Assess decision trees' ability to recover cluster structuresInvestigate tighter guarantees for well-clustered dataStudy explainable clustering beyond worst-case guarantees

Latest Papers

What's happening recently
View more

This work addresses the lack of rigorous mathematical definitions for “concepts” and “learning” in existing sparse autoencoders, which obscures the mechanisms underlying neuron interpretability. The authors formalize concepts as sets of data points and frame concept learning as a set-alignment problem between human-defined concepts and model-induced concepts, distinguishing three hierarchical levels: detection, separation, and approximation. Building on geometric and set-theoretic foundations, they propose a unified theoretical framework that elucidates the origins of phenomena such as feature splitting, absorption, familial relationships, and hierarchical structure. Leveraging formal concept analysis, the framework captures the many-to-many correspondence between neurons and concepts. Theoretical predictions are validated on synthetic data, revealing how model scale and sparsity jointly influence concept learning capacity, and enabling the construction of concept lattices that systematically organize neuron–concept mappings.

concept learningfeature representationgeometric understanding

Existing mechanistic interpretability methods typically focus on individual prompt–output pairs, making it difficult to uncover the underlying heterogeneity of mechanisms across a language model’s generation distribution. This work proposes an unsupervised feature discovery approach that clusters model-generated continuations by jointly leveraging semantic embeddings and attribution signatures from prefix-to-continuation mappings. Without requiring human-specified target outputs, this method achieves, for the first time, mechanism–semantic alignment at the distributional level. It optimizes a rate–distortion objective that balances semantic coherence, mechanistic consistency, and cluster granularity, effectively revealing diverse continuation mechanisms invisible to single-perspective analyses. Intervention experiments further validate that the learned cluster signatures correspond to manipulable internal computational factors, substantially enhancing the scalability of model auditing.

attribution signaturescircuit analysiscontinuation distribution

This work addresses the limited interpretability utility of sparse autoencoder (SAE) features due to their unstable causal influence on model behavior. To systematically analyze the downstream effects of feature interventions, the authors propose the Feature Effect Geometric Analysis (FEGA) framework, which—through cross-context ablation of identical SAE features and modeling of their effect geometry from the perspective of output logit changes—reveals that most SAE features lack consistent one-dimensional effects. Integrating unsupervised ablation, geometric analysis, and feature categorization, FEGA demonstrates that value-like features exhibit low-dimensional yet multidirectional impacts, whereas pointer-like features display more diffuse effects. The study further shows that interpretability does not necessarily entail stable, controllable directions of influence.

causal effectsfeature effectsinterpretability

This work addresses a critical limitation in existing neuron-level concept explanation methods, which often assume that all neurons possess clear functional roles, thereby overlooking redundant or misleading neurons that can distort interpretations of model decision-making. To overcome this, the authors propose the Select-Hypothesize-Verify (SHV) framework: it first selects the most representative samples based on activation distributions, then generates natural language concept hypotheses, and finally validates these hypotheses through a neuron activation verification mechanism. SHV introduces, for the first time, a systematic pipeline for concept validation, effectively identifying and focusing on neurons with genuine semantic meaning. Experimental results demonstrate that concepts produced by SHV activate target neurons at 1.5 times the rate of state-of-the-art methods, substantially improving the accuracy and reliability of model interpretations.

concept verificationmisleading neuron conceptsneural network interpretability

Hot Scholars

GN

Graham Neubig

Carnegie Mellon University, All Hands AI
Natural Language ProcessingMachine LearningArtificial Intelligence
SK

Seungone Kim

Carnegie Mellon University
Large Language ModelsNatural Language Processing
YL

Yadan Luo

ARC DECRA and Senior Lecturer, University of Queensland
Generalization3D VisionAutonomous Driving
HD

Henghui Ding

Fudan University
Computer VisionMachine LearningSegmentationAIGC