Quantifier Scope Interpretation in Language Learners and LLMs

📅 2025-09-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates large language models’ (LLMs’) preferences in quantifier scope interpretation within English and Chinese sentences containing multiple quantifiers, and their alignment with human cognitive judgments. We propose a probabilistic interpretation evaluation framework, calibrated against cross-linguistic human experimental data, and introduce the Human Similarity (HS) score to quantify model–human agreement in scope judgments. Results show that most LLMs exhibit a strong preference for surface-scope interpretations; certain large-scale models capture the cross-linguistic reversal in scope preferences—namely, subject-wide scope dominance in English versus object-wide scope dominance in Chinese. Model scale, architecture, and proportion of Chinese pretraining data significantly influence scope modeling capability. This work constitutes the first systematic examination of LLMs’ cross-linguistic quantifier scope reasoning at the formal semantic level, establishing an interpretable, comparable paradigm for evaluating deep linguistic understanding in foundation models.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Learning Human Values and Preferences

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Sentences with multiple quantifiers often lead to interpretive ambiguities, which can vary across languages. This study adopts a cross-linguistic approach to examine how large language models (LLMs) handle quantifier scope interpretation in English and Chinese, using probabilities to assess interpretive likelihood. Human similarity (HS) scores were used to quantify the extent to which LLMs emulate human performance across language groups. Results reveal that most LLMs prefer the surface scope interpretations, aligning with human tendencies, while only some differentiate between English and Chinese in the inverse scope preferences, reflecting human-similar patterns. HS scores highlight variability in LLMs' approximation of human behavior, but their overall potential to align with humans is notable. Differences in model architecture, scale, and particularly models' pre-training data language background, significantly influence how closely LLMs approximate human quantifier scope interpretations.
Problem

Research questions and friction points this paper is trying to address.

Quantifier scope interpretation ambiguities across languages
How LLMs handle quantifier scope in English and Chinese
Extent LLMs emulate human performance in scope interpretation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-linguistic probabilistic assessment method
Human similarity scoring for model evaluation
Analysis of architecture and training data effects
🔎 Similar Papers
No similar papers found.
S
Shaohua Fang
Department of English, Purdue University
Y
Yue Li
Department of Linguistics, Purdue University
Yan Cong
Yan Cong
Purdue University
semanticspragmaticsChinese linguisticsnatural language processingclinical linguistics