🤖 AI Summary
This study addresses the prevalent conflation of feature manipulation effects with intrinsic computational mechanisms in large language model (LLM) research, wherein behavioral changes are erroneously taken as evidence of actual feature utilization. To resolve this, we propose an empirical contract grounded in natural input values, employing causal intervention techniques—including activation copying, removal, and downstream recovery—to conduct comparative experiments across diverse LLM representations. This framework effectively disentangles manipulation strength from the degree of model reliance. Our findings reveal that no single metric suffices to demonstrate genuine feature utilization and uncover significant disparities in the extent to which different representations can be manipulated versus utilized. Ultimately, this work corrects the cognitive bias of inferring internal mechanisms solely from external intervention outcomes, offering a more rigorous methodological foundation for mechanistic interpretability in LLMs.
📝 Abstract
Internal features in LLMs are often interpreted as mechanisms when they track a concept and their manipulation changes a related behavior. Yet steering can push a feature far outside its natural range, where its effects need not reflect the model's own computation. We examine this inference and propose an empirical contract whose tests evaluate features at values observed on natural inputs. One test copies a feature's value from an input that shows a behavior into a matched input that does not (installation) or the reverse (removal); the other restores the feature after an upstream edit (downstream rescue). Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model uses it. Applied to three kinds of representations, the two strengths separate sharply. The published unknown-entity latent strongly steers knowledge abstention, yet installing observed values from either published latent into matched prompts transfers only a small fraction of the natural known--unknown abstention contrast. Dense known--unknown directions show opposite asymmetries between installation and removal in Gemma and Llama, and how fully a released subject--verb agreement feature set reproduces and restores the behavior depends on how its values are written into the model. Tracking a concept and steering a behavior therefore do not by themselves show that the model uses a feature, and each conclusion holds only for the intervention tested.