Which Attention Heads are like the Human Head? Not the Ones that Compute

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether brain-AI alignment is equivalent to computational causality in models. By comparing attention heads of large language models with human EEG signals and integrating attribution patching with concept and function vector analyses, the authors conduct ablation experiments across 17 models to identify two distinct categories of attention heads: novelty and repetition. The findings demonstrate that brain alignment reflects input reading rather than task solving. Specifically, removing components based on brain alignment yields a substantially smaller performance degradation than attribution-based removal, revealing a significant decoupling between the two. This work provides critical empirical evidence for understanding the computational nature of neural alignment.
📝 Abstract
Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA $\rightarrow$ B), we compare LLM attention-head representations with human EEG and test how ablating those heads affects task performance. Alignment and causation dissociate: brain-aligned heads contribute to performance, but their removal is substantially less disruptive than removal of heads selected via attribution patching. We compare two head sets that prior interpretability work defines without reference to the brain: concept vectors (CVs), which represent abstract patterns across formats, and function vectors (FVs), selected for their contribution to correct-answer prediction. Brain alignment shows little association with FV scores, while its association with CV scores varies across models. Among brain-aligned heads, we find recurring attention profiles: one emphasizes distinctive elements (novelty heads), the other repeating elements (repetition heads). The novelty family tracks salience and attends to the same elements that humans look at, yet its removal is less damaging than random ablation on average. Repetition heads contribute modestly to performance and are associated with abstract-pattern representation (CVs). Across 17 models spanning 3B-72B parameters, FV-ranked removal is substantially more disruptive than brain-ranked removal. Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task.
Problem

Research questions and friction points this paper is trying to address.

Brain-AI alignment
causal contribution
attention heads
large language models
EEG
Innovation

Methods, ideas, or system contributions that make the work stand out.

Brain-AI alignment
Attention heads ablation
Causal dissociation
Concept vectors
Function vectors
🔎 Similar Papers
No similar papers found.