🤖 AI Summary
This work investigates the internal mechanisms by which large language models (LLMs) perform relevance judgment in information retrieval. Addressing the open question—“How do LLMs understand and model query-document relevance?”—we propose the first mechanism-based, multi-stage interpretability framework. Our analysis reveals that early layers extract semantic features, intermediate layers activate relevance-specific reasoning pathways conditioned on instructions, and late-layer attention heads generate structured relevance judgments. Methodologically, we integrate activation patching with layer- and head-level attribution analysis to precisely localize functional roles across model components. Empirical results demonstrate that LLMs possess explicit, stage-wise relevance modeling capabilities—not merely opaque matching. This study uncovers an interpretable cognitive architecture underlying LLM-based IR and provides both theoretical foundations and design principles for developing trustworthy, controllable LLM-driven retrieval systems.
📝 Abstract
Recent studies have shown that large language models (LLMs) can assess relevance and support information retrieval (IR) tasks such as document ranking and relevance judgment generation. However, the internal mechanisms by which off-the-shelf LLMs understand and operationalize relevance remain largely unexplored. In this paper, we systematically investigate how different LLM modules contribute to relevance judgment through the lens of mechanistic interpretability. Using activation patching techniques, we analyze the roles of various model components and identify a multi-stage, progressive process in generating either pointwise or pairwise relevance judgment. Specifically, LLMs first extract query and document information in the early layers, then process relevance information according to instructions in the middle layers, and finally utilize specific attention heads in the later layers to generate relevance judgments in the required format. Our findings provide insights into the mechanisms underlying relevance assessment in LLMs, offering valuable implications for future research on leveraging LLMs for IR tasks.