mechanistic interpretability

Designs and builds circuit‑level maps and analyses of a model’s internal mechanisms by identifying and characterizing components (neurons, attention heads, layers, pathways), tracing activations across layers, and locating architectural roles or bottlenecks. Implements and evaluates causal probes and interventions to attribute behaviors or failures to specific internal structures, compare intervention strategies, and validate mechanistic explanations.

mechanisticinterpretability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$240K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether internal circuits in language models exhibit task-specificity and consistency, and how such properties inform our understanding of—and ability to intervene on—model behavior. Employing edge attribution patching and component ablation, the authors systematically evaluate causally critical subgraphs within attention heads and MLP layers across six tasks and seven models. Their analysis reveals, for the first time, that circuits within a single task are highly reused and essential for performance, yet circuits across different tasks substantially overlap, with task-exclusive components contributing minimally. This finding challenges the prevailing assumption of task-dedicated circuits and offers a new perspective on model interpretability and targeted intervention.

circuitsconsistencylanguage models

Uncovering Intermediate Variables in Transformers using Circuit Probing

Nov 07, 2023
MA
Michael A. Lepori
🏛️ Brown University

Understanding the causal roles and computational mechanisms of intermediate variables—such as syntactic attributes—in Transformer language models remains challenging. Method: We propose *circuit probing*, a hypothesis-driven methodology that reverse-engineers intermediate representations encoding specific linguistic properties and precisely identifies the parameter-level neural circuits supporting them. Our approach integrates gradient-guided discovery, targeted parameter ablation, verification via diagnostic probe training, and modular attribution analysis to enable causal intervention and algorithm-level interpretation. Contribution/Results: This work unifies hypothesis testing, circuit localization, and dynamic tracing for the first time, revealing implicit algorithmic structures within models and their training-time evolution. Experiments successfully decode symbolic arithmetic logic in specialized arithmetic models, localize subject–verb agreement and reflexive pronoun processing circuits in GPT-2, and empirically confirm their progressive emergence during training.

Identify intermediate computation variablesInterpret neural network algorithmsTest syntactic property hypotheses

Existing circuit analysis methods are fragmented and lack a unified framework to support the end-to-end pipeline from discovery and evaluation to downstream interventions, often relying on manually crafted contrastive prompts that hinder reproducibility and scalability. This work proposes the first end-to-end circuit analysis toolkit, built upon a typed, serializable circuit representation that integrates multiple discovery algorithms, declarative task mapping, diagnostic utilities, and intervention modules—including pruning, editing, and steering—to enable fully interpretable mechanistic analysis throughout the pipeline. The framework facilitates automated contrastive prompt generation, cross-task transfer, and algorithmic comparison, substantially lowering barriers to both research and practical application. The complete library, along with examples and documentation, has been open-sourced to provide the community with standardized, reusable infrastructure.

circuit analysiscontrastive promptsdownstream interventions

Existing interpretability methods for vision models predominantly focus on neuron activations, lacking a mechanistic understanding of how information propagates through the model. This work proposes Visual Circuit Discovery (Vi-CD), the first method to introduce edge-level mechanistic circuit analysis into Vision Transformers. By constructing an edge-based computational graph, Vi-CD automatically identifies class-specific information pathways relevant to particular tasks. The approach not only uncovers the internal information routing mechanisms of vision models but also successfully locates adversarial circuits in CLIP that underlie typographic attacks. Furthermore, targeted interventions on these discovered circuits effectively mitigate harmful behaviors, enabling transparent and controllable manipulation of visual model decision processes.

computational graphsedge-based circuitsmechanistic interpretability

The causal origins of interpretable units—such as induction heads—in large language models remain poorly understood. This work proposes a scalable mechanistic data attribution framework that integrates influence functions with causal interventions to establish, for the first time, direct causal links between specific training examples and the emergence of such interpretable components. The study reveals that structured repetitive data plays a catalytic role in circuit formation and demonstrates a direct functional relationship between induction heads and in-context learning capabilities. By selectively intervening on a small set of high-influence training samples, the emergence of attention heads can be significantly modulated. Furthermore, the proposed data augmentation strategy consistently accelerates circuit convergence across different model scales.

Data AttributionIn-Context LearningInduction Heads

Latest Papers

What's happening recently
View more

This work addresses the challenge in mechanistic interpretability that, despite progress in circuit localization, component-level functional explanations remain manual and lack standardization. To this end, we propose HyVE, a novel framework that introduces language model agents into circuit explanation tasks. HyVE iteratively performs observation, hypothesis generation, and causal verification to automatically produce both component-level interpretations and circuit-level task descriptions. We construct AgenticInterpBench, the first benchmark tailored for agent-based interpretability, and evaluate HyVE across four mainstream language model architectures, demonstrating its ability to generate high-quality explanations. Our experiments reveal that causal verification constitutes the primary performance bottleneck, and we further showcase HyVE’s practical utility through a case study on arithmetic circuits in Llama-3-8B.

circuit explanationcomponent-level explanationlanguage model agents

This work addresses the challenge of ensuring safety and auditability in high-stakes applications of modern neural networks, which are often hindered by their opaque “black-box” nature. The authors propose a mechanistic interpretability framework that integrates Transformer circuit analysis, sparse autoencoders (SAEs), and neuro-symbolic reasoning to decompose complex activations into human-interpretable sparse features. This approach identifies critical computational units—such as induction heads—and establishes a mapping from neural representations to executable logical rules. By enabling causal interventions and targeted control over model behavior, the method achieves end-to-end symbolic explanations and controllability for the first time, significantly enhancing model transparency and intervenability without compromising performance.

black boxmechanistic interpretabilityneural networks

Hot Scholars

MG

Mor Geva

Tel Aviv University, Google Research
Natural Language Processing
GB

Gemma Boleda

ICREA Research Professor, Universitat Pompeu Fabra
linguisticscomputational linguisticscognitive sciencesemantics
VK

Vinay Kumar Sankarapu

CEO & Founder, Arya.ai & AryaXAI.com
Deep LearningAI InterpretabilityAI AlignmentReinforcement Learning
PS

Pratinav Seth

AryaXAI Alignment Lab, Arya.ai (An Aurionpro Company)
Deep LearningExplainable AIAI for RiskAI for Social Good
AG

Atticus Geiger

Pr(Ai)²R Group
Artificial IntelligenceNatural LanguageMechanistic InterpretabilityCausality