Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates how language models internally represent network infrastructure information, such as hostname-IP pairs. Through causal ablation and intervention experiments, the authors localize and validate specific attention heads and neurons responsible for this task, evaluating their generalizability via correlation-based ranking and cross-dataset transfer tests. The findings demonstrate that attention-head-level causal mechanisms exhibit cross-architecture universality, with full-head detectors achieving 99.5%โ€“100% accuracy across multiple models. Conversely, neuron-level responsibility distributions are shown to be model-specific, limiting the transferability of single-neuron approaches and necessitating per-model validation.
๐Ÿ“ Abstract
We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100\% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100\% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.
Problem

Research questions and friction points this paper is trying to address.

Language Models
Attention Heads
Neurons
Network Information Retrieval
Causal Validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal ablation
attention heads
mechanistic interpretability
single neuron
network information retrieval
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.