Score
Designs and implements unsupervised methods to discover and characterize interpretable features in neural network activations by mining activation geometry (directions, clusters, manifolds) and extracting feature vectors or readout directions without labeled supervision. Builds analyses and procedures to approximate features as activation directions, quantify how those features change or influence downstream readouts, and produce human‑interpretable reasoning features derived from activations.
This work addresses the limitations of existing interpretability methods that rely on human-annotated concepts, which are prone to bias and struggle to uncover reasoning features in large language models in an unsupervised manner. The authors propose MAG, an unsupervised framework that prepends a unified natural language instruction to inputs and measures its impact on the model’s activation space to automatically extract reasoning-relevant semantic features—without requiring labeled data. They find that certain features can be approximated by single activation directions, enabling intervention via vector arithmetic. Through activation geometry analysis, activation differencing, vector manipulation, and RFD similarity metrics, the extracted features accurately predict the model’s world knowledge. In prompt injection classification tasks, selecting training data using RFD yields 94.7% Top-1 and 100% Top-2 accuracy.
This work addresses the lack of human-interpretable concepts in intermediate-layer representations of CNNs. We propose an unsupervised post-hoc method that optimizes an orthogonal rotation in feature space to extract disentangled, concept-level interpretable basis vectors from sparsely thresholded activation responses. Unlike supervised approaches relying on manually annotated concepts, ours is the first purely unsupervised paradigm for discovering highly interpretable bases. We further introduce an improved interpretability metric and a concept-alignment analysis framework, validating our method across multiple CNN architectures and datasets. Experiments demonstrate that the rotated intermediate representations significantly outperform supervised basis extraction methods in both conceptual diversity and interpretability. Our results reveal an inherent limitation of supervised paradigms—namely, their restricted coverage of conceptual breadth—and open a new direction for model interpretability research. (149 words)
To address challenges in unsupervised knowledge discovery—including weak modeling of feature correlations, poor pattern interpretability, and semantic distortion in dimensionality reduction—this paper proposes a three-stage analytical framework grounded in an Unsupervised Cognition model: (1) association pattern mining, (2) cognition-significance-driven interpretable feature selection, and (3) semantic-consistency-constrained dimensionality reduction. It pioneers the end-to-end integration of cognitive modeling with knowledge discovery, enabling fully label-free, interpretable analysis. Evaluated on diverse multi-source empirical datasets, the method consistently outperforms state-of-the-art approaches across all core metrics: pattern completeness (+12.7%), feature discriminability (+9.4%), and dimensionality-reduction interpretability (+15.3%).
This paper addresses the challenge of identifying hierarchical, distributed activation patterns in neural networks. To overcome limitations of neuron-level or hand-crafted interpretable feature analyses, we propose Neural Activation Pattern (NAP) modeling based on full-layer activation distributions. Methodologically, we introduce normalized preprocessing, kernel density estimation for distribution modeling, SNR-adaptive distance metrics, and hierarchical clustering. For the first time, our approach uncovers a continuous activation manifold in neural communication receivers that is dominantly governed by signal-to-noise ratio (SNR), demonstrating its physical interpretability. Experiments show that NAP significantly improves separation between in-distribution and out-of-distribution samples, empirically confirming SNR as a critical implicit factor learned by the model. The method enhances generalization, interpretability, and reliability diagnostics—enabling robust, physics-informed analysis of deep neural representations.
This work addresses a fundamental challenge in interpretable clustering: whether the worst-case interpretability cost bound can be surpassed—and the underlying cluster structure reliably recovered—when data exhibit well-separated clusters. To this end, we propose a decision-tree-based mixture-model clustering method. We introduce the first theoretical framework of “interpretability–noise ratio,” enabling data-agnostic, efficient tree construction. Under sub-Gaussian assumptions, we derive tight upper and lower bounds on estimation error. Furthermore, we pioneer the integration of Concept Activation Vectors (CAVs) into unsupervised clustering, facilitating interpretable cluster identification in deep representation spaces. Experiments on standard tabular and image benchmarks demonstrate that our method significantly enhances interpretability while maintaining high clustering accuracy; moreover, its theoretical guarantees strictly improve upon those of existing distribution-agnostic approaches.
该论文探讨了神经网络中的特征叠加问题,通过理论和实践方法分析了特征叠加的几何、学习和计算,并评估了从训练网络中恢复和分析特征的方法。
This work addresses the lack of rigorous mathematical definitions for “concepts” and “learning” in existing sparse autoencoders, which obscures the mechanisms underlying neuron interpretability. The authors formalize concepts as sets of data points and frame concept learning as a set-alignment problem between human-defined concepts and model-induced concepts, distinguishing three hierarchical levels: detection, separation, and approximation. Building on geometric and set-theoretic foundations, they propose a unified theoretical framework that elucidates the origins of phenomena such as feature splitting, absorption, familial relationships, and hierarchical structure. Leveraging formal concept analysis, the framework captures the many-to-many correspondence between neurons and concepts. Theoretical predictions are validated on synthetic data, revealing how model scale and sparsity jointly influence concept learning capacity, and enabling the construction of concept lattices that systematically organize neuron–concept mappings.
Existing mechanistic interpretability methods typically focus on individual prompt–output pairs, making it difficult to uncover the underlying heterogeneity of mechanisms across a language model’s generation distribution. This work proposes an unsupervised feature discovery approach that clusters model-generated continuations by jointly leveraging semantic embeddings and attribution signatures from prefix-to-continuation mappings. Without requiring human-specified target outputs, this method achieves, for the first time, mechanism–semantic alignment at the distributional level. It optimizes a rate–distortion objective that balances semantic coherence, mechanistic consistency, and cluster granularity, effectively revealing diverse continuation mechanisms invisible to single-perspective analyses. Intervention experiments further validate that the learned cluster signatures correspond to manipulable internal computational factors, substantially enhancing the scalability of model auditing.
This work addresses the limited interpretability utility of sparse autoencoder (SAE) features due to their unstable causal influence on model behavior. To systematically analyze the downstream effects of feature interventions, the authors propose the Feature Effect Geometric Analysis (FEGA) framework, which—through cross-context ablation of identical SAE features and modeling of their effect geometry from the perspective of output logit changes—reveals that most SAE features lack consistent one-dimensional effects. Integrating unsupervised ablation, geometric analysis, and feature categorization, FEGA demonstrates that value-like features exhibit low-dimensional yet multidirectional impacts, whereas pointer-like features display more diffuse effects. The study further shows that interpretability does not necessarily entail stable, controllable directions of influence.
This work addresses a critical limitation in existing neuron-level concept explanation methods, which often assume that all neurons possess clear functional roles, thereby overlooking redundant or misleading neurons that can distort interpretations of model decision-making. To overcome this, the authors propose the Select-Hypothesize-Verify (SHV) framework: it first selects the most representative samples based on activation distributions, then generates natural language concept hypotheses, and finally validates these hypotheses through a neuron activation verification mechanism. SHV introduces, for the first time, a systematic pipeline for concept validation, effectively identifying and focusing on neurons with genuine semantic meaning. Experimental results demonstrate that concepts produced by SHV activate target neurons at 1.5 times the rate of state-of-the-art methods, substantially improving the accuracy and reliability of model interpretations.