🤖 AI Summary
This study investigates neural correspondences between large language models (LLMs) and human cortical language processing. To bridge computational and neural levels, we employ explainable AI (XAI) attribution methods—specifically Integrated Gradients and occlusion analysis—to quantify token-level predictive contributions of LLMs and directly map these attributions to fMRI-measured brain activity during naturalistic narrative comprehension. Our key contributions are threefold: (1) Attribution maps—unlike standard hidden-layer representations—more accurately capture the hierarchical organization of cortical language processing; (2) attribution magnitude exhibits significant positive correlation with neural response strength, empirically validating alignment between LLM computational depth and functional neuroanatomy; and (3) we introduce a brain-alignment-based evaluation framework to assess the biological plausibility of attribution methods. Results demonstrate that attribution-based predictions significantly outperform conventional representation-based baselines across the whole language network, particularly in mid- and early-stage language regions.
📝 Abstract
Recent advances in artificial intelligence have given rise to large language models (LLMs) that not only achieve human-like performance but also share computational principles with the brain's language processing mechanisms. While previous research has primarily focused on aligning LLMs' internal representations with neural activity, we introduce a novel approach that leverages explainable AI (XAI) methods to forge deeper connections between the two domains. Using attribution methods, we quantified how preceding words contribute to an LLM's next-word predictions and employed these explanations to predict fMRI recordings from participants listening to the same narratives. Our findings demonstrate that attribution methods robustly predict brain activity across the language network, surpassing traditional internal representations in early language areas. This alignment is hierarchical: early-layer explanations correspond to the initial stages of language processing in the brain, while later layers align with more advanced stages. Moreover, the layers more influential on LLM next-word prediction$unicode{x2014}$those with higher attribution scores$unicode{x2014}$exhibited stronger alignment with neural activity. This work establishes a bidirectional bridge between AI and neuroscience. First, we demonstrate that attribution methods offer a powerful lens for investigating the neural mechanisms of language comprehension, revealing how meaning emerges from preceding context. Second, we propose using brain alignment as a metric to evaluate the validity of attribution methods, providing a framework for assessing their biological plausibility.