Grammatical "grandmother neurons" are rare in LLMs

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional probing methods in disentangling intrinsic grammatical representations from classifier capabilities in large language models, a challenge arising from capacity confounds and calibration biases. To overcome this, we propose a probeless framework that eliminates auxiliary classifier training. Leveraging linguistic minimal pairs and a Neuron Separability Index (NSI) combined with permutation normalization, our approach directly quantifies single-neuron grammatical selectivity without parameter updates, while causal dependencies are validated through targeted ablation. Our findings reveal that morphosyntactic dissociation emerges prior to semantic interfaces, and that highly selective “grandmother neurons” are exceedingly rare. Crucially, we demonstrate that neuronal activation selectivity is significantly decoupled from both model behavioral performance and causal reliance.
📝 Abstract
Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations, we introduce a probe-free framework for localizing linguistic selectivity at the individual neuron level. Leveraging the controlled contrasts of linguistic minimal pairs, we propose a Neuron Separability Index (NSI), a metric that directly quantifies how reliably single neurons differentiate grammatical from ungrammatical constructions without parameter updates. Applying NSI across 68 linguistic paradigms and seven checkpoints reveals three main patterns: 1) raw separability reaches near-peak levels earlier for morphological and syntactic distinctions than for syntax-semantics interface and conceptual distinctions. 2) after permutation normalization, single-unit selectivity is sparse, weak, and narrowly tuned: only a small fraction of units are sensitive to an average paradigm, and strongly selective "grandmother neurons" are rare. 3) whole-vector linear separability, single-neuron selectivity, and behavioral competence are largely dissociated, and targeted ablations further separate activation selectivity from causal reliance.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
interpretability
linguistic representations
grandmother neurons
probe-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

probe-free framework
Neuron Separability Index (NSI)
linguistic minimal pairs
grandmother neurons
interpretability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Linyang He
Zuckerman Mind Brain Behavior Institute, Columbia University
Nima Mesgarani
Nima Mesgarani
Associate Professor, Columbia University
Speech neurosciencespeech modelingspeech technologies