Inferring Functionality of Attention Heads from their Parameters

๐Ÿ“… 2024-12-16
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 3
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of interpreting attention head functionality in large language models (LLMs), where conventional methods rely on costly inference or training. We propose MAPS, a parameter-only framework that infers head-level functional semantics without forward passesโ€”leveraging geometric analysis in parameter space, decomposition of attention weights, functional template matching, and causal intervention validation. MAPS is the first method to enable purely parameter-driven functional mapping of attention heads, supporting quantitative measurement of operational strength and identification of individual head functionality. It uncovers previously overlooked functional patterns, revealing both functional universality across models and architecture-specific biases. Evaluated on six mainstream LLMs, MAPS achieves high correlation (>0.85) for 20 distinct operations; human evaluation confirms >90% of its functional descriptions as semantically reasonable. Furthermore, it systematically characterizes cross-model functional distribution patterns.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
Attention heads are one of the building blocks of large language models (LLMs). Prior work on investigating their operation mostly focused on analyzing their behavior during inference for specific circuits or tasks. In this work, we seek a comprehensive mapping of the operations they implement in a model. We propose MAPS (Mapping Attention head ParameterS), an efficient framework that infers the functionality of attention heads from their parameters, without any model training or inference. We showcase the utility of MAPS for answering two types of questions: (a) given a predefined operation, mapping how strongly heads across the model implement it, and (b) given an attention head, inferring its salient functionality. Evaluating MAPS on 20 operations across 6 popular LLMs shows its estimations correlate with the head's outputs during inference and are causally linked to the model's predictions. Moreover, its mappings reveal attention heads of certain operations that were overlooked in previous studies, and valuable insights on function universality and architecture biases in LLMs. Next, we present an automatic pipeline and analysis that leverage MAPS to characterize the salient operations of a given head. Our pipeline produces plausible operation descriptions for most heads, as assessed by human judgment, while revealing diverse operations.
Problem

Research questions and friction points this paper is trying to address.

Infer functionality of attention heads from parameters
Map predefined operations across model heads
Automatically characterize salient head operations
Innovation

Methods, ideas, or system contributions that make the work stand out.

MAPS infers head functionality from parameters
No model training or inference required
Automatically characterizes head operations
๐Ÿ”Ž Similar Papers
No similar papers found.