Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
This study investigates how language models internally represent network infrastructure information, such as hostname-IP pairs. Through causal ablation and intervention experiments, the authors localize and validate specific attention heads and neurons responsible for this task, evaluating their generalizability via correlation-based ranking and cross-dataset transfer tests. The findings demonstrate that attention-head-level causal mechanisms exhibit cross-architecture universality, with full-head detectors achieving 99.5%–100% accuracy across multiple models. Conversely, neuron-level responsibility distributions are shown to be model-specific, limiting the transferability of single-neuron approaches and necessitating per-model validation.