🤖 AI Summary
This study investigates whether explicitly exposing routing mechanisms in Transformers is sufficient to achieve mechanistic interpretability. To this end, the authors propose Block Attention Residuals, which represent cross-layer information routing as observable tensors during forward propagation and enable causal intervention to analyze their functional roles. Experiments based on the Qwen3 architecture demonstrate that meaningful local routing patterns emerge only when the routing structure is actively involved in training optimization. Crucially, the magnitude of routing weights does not directly reflect causal importance, necessitating intervention-based validation of interpretability hypotheses. The work identifies three characteristic local routing patterns and reveals that segments with the largest routing weights do not necessarily contribute the most causally, establishing that explicit exposure of routing mechanisms is necessary but insufficient for mechanistic interpretability.
📝 Abstract
Block Attention Residuals (Block AttnRes) by replace fixed additive residuals with a learned softmax over earlier depth-source representations, surfacing cross-layer routing as an inspectable tensor in the forward pass. This is a tempting interpretability target: information flow normally inferred indirectly is now directly observable. We ask whether such exposure suffices for mechanistic interpretation. We probe two same-scale ($0.6$B) Block AttnRes checkpoints under identical routing-ablation interventions: a vanilla Qwen3 inference-wrapped through a deterministic recency-bias schedule that the codebase admits as a routing-equivalent loading path, and a Block AttnRes Qwen3 trained from scratch with routing as part of optimisation. The wrapped baseline's routing weights are content-independent and reproduce the schedule's analytic prediction. The trained AttnRes checkpoint instead exhibits three localised routing motifs: an embedding-source pathway through early-layer MLP, a current-state pathway through early-layer attention and MLP, and an older-history pathway through late-layer attention. Beyond this stratification, we find a sharp dissociation between average routing mass and causal importance: in both sublayers, the largest mass slice is not the largest causal contribution, and one source family carries appreciable mass with no detectable causal role under intervention. Architectural exposure of routing is therefore necessary but not sufficient for mechanistic interpretation: structured depth routing emerges only when routing has been part of training, and even then, descriptive routing summaries should be treated as candidate hypotheses to be tested by causal interventions, not as evidence of mechanism in their own right.