How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective

📅 2025-04-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the internal mechanisms by which large language models (LLMs) perform relevance judgment in information retrieval. Addressing the open question—“How do LLMs understand and model query-document relevance?”—we propose the first mechanism-based, multi-stage interpretability framework. Our analysis reveals that early layers extract semantic features, intermediate layers activate relevance-specific reasoning pathways conditioned on instructions, and late-layer attention heads generate structured relevance judgments. Methodologically, we integrate activation patching with layer- and head-level attribution analysis to precisely localize functional roles across model components. Empirical results demonstrate that LLMs possess explicit, stage-wise relevance modeling capabilities—not merely opaque matching. This study uncovers an interpretable cognitive architecture underlying LLM-based IR and provides both theoretical foundations and design principles for developing trustworthy, controllable LLM-driven retrieval systems.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Recent studies have shown that large language models (LLMs) can assess relevance and support information retrieval (IR) tasks such as document ranking and relevance judgment generation. However, the internal mechanisms by which off-the-shelf LLMs understand and operationalize relevance remain largely unexplored. In this paper, we systematically investigate how different LLM modules contribute to relevance judgment through the lens of mechanistic interpretability. Using activation patching techniques, we analyze the roles of various model components and identify a multi-stage, progressive process in generating either pointwise or pairwise relevance judgment. Specifically, LLMs first extract query and document information in the early layers, then process relevance information according to instructions in the middle layers, and finally utilize specific attention heads in the later layers to generate relevance judgments in the required format. Our findings provide insights into the mechanisms underlying relevance assessment in LLMs, offering valuable implications for future research on leveraging LLMs for IR tasks.
Problem

Research questions and friction points this paper is trying to address.

How LLMs internally assess relevance for IR tasks
Roles of different LLM modules in relevance judgment
Mechanistic process of generating relevance judgments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation patching analyzes LLM module roles
Multi-stage process generates relevance judgments
Specific attention heads format final outputs
🔎 Similar Papers
2024-05-03Annual International ACM SIGIR Conference on Research and Development in Information RetrievalCitations: 2