🤖 AI Summary
This study investigates whether large language models (LLMs) exhibit human-like anaphora resolution capabilities, particularly under the influence of discourse structure, situation models, and semantic factors. Leveraging surprisal and comprehension question accuracy as evaluation metrics, the research systematically assesses the sensitivity of five open-source LLMs to discourse prominence, distance effects, and semantic interference. Grounded in the standard linking hypothesis from cognitive science, model surprisal is correlated with human reading times and complemented by behavioral experiments for comparative analysis. The findings indicate that certain models successfully replicate human sensitivity to prominence and distance effects; however, they generally fail to respond appropriately to semantic interference, revealing a critical limitation in current LLMs’ ability to simulate human-like mechanisms of anaphoric processing.
📝 Abstract
Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.