From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

📅 2026-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of moving beyond correlational analyses to establish causal relationships in identifying truly influential features within Transformer models. It proposes a five-stage causal analysis framework—comprising probe design, feature extraction, causal validation, robustness testing, and deployment integration—to systematically evaluate features in GPT-2 small on the Indirect Object Identification (IOI) task. By integrating activation patching, sparse autoencoders, feature ablation, and an NLA-inspired variance attribution method, the study reveals a negative correlation between feature selectivity and causal efficacy, and uncovers a substantial gap between detection robustness and causal robustness. While replicating the IOI circuit, it finds that only a subset of 15 highly selective features exhibit genuine causal effects, and model performance degrades significantly under distributional shift. A cost-aware monitoring strategy achieves 99.1% resource savings.
📝 Abstract
We propose a five-stage methodology for causal feature analysis in transformer language models (probe design, feature extraction, causal validation, robustness testing, and deployment integration) and demonstrate it end-to-end on GPT-2 small performing the Indirect Object Identification (IOI) task. Activation patching recovers the canonical IOI circuit (layer-9 head 9 alone gives recovery +1.02). A sparse autoencoder recovers per-name selective features with effect sizes of 30 to 50 activation units. Causal validation finds these features specifically but only partially causal: ablating fifteen of them leaves the model accurate on 98% of prompts. Two NLA-inspired evaluations strengthen this picture: the fifteen selective features explain only 31% of activation variance versus the SAE's 99.7%, and selectivity ratio anticorrelates with causal force (r = -0.56). Robustness testing under three distribution shifts finds that the circuit transfers cleanly but feature ablation effects degrade substantially, exposing a gap between detection robustness and causal robustness. A cost-based deployment evaluation (assumed $50/FN, $0.42/FP, 2% error rate) finds an optimal monitor configuration yielding $8.96 per 1000 queries against a $1000 baseline, a 99.1% saving. Optimal composition strategy varies with cost ratio and base rate. The conjunction of stages produces findings no single stage would.
Problem

Research questions and friction points this paper is trying to address.

causal feature analysis
transformer language models
Indirect Object Identification
feature causality
correlation vs. causation
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal feature analysis
sparse autoencoder
activation patching
robustness testing
cost-based deployment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Caleb Munigety
Independent Researcher