Investigating Syntactic Biases in Multilingual Transformers with RC Attachment Ambiguities in Italian and English

📅 2025-04-14
📈 Citations: 0
Influential: 0
📄 PDF

career value

147K/year
🤖 AI Summary
This study investigates whether monolingual and multilingual large language models (LLMs) replicate human syntactic preferences in relative clause (RC) attachment ambiguity tasks in Italian and English, and to what extent lexical factors—such as main-clause verb/noun type—modulate these preferences. Using a psycholinguistically grounded, controlled cross-lingual stimulus set, the authors employ probabilistic output analysis, lexical embedding perturbation, and cross-lingual comparative evaluation. Results reveal that mainstream LLMs consistently lack human-like hierarchical attachment biases and exhibit weak cross-lingual consistency; lexical factors—including verb semantic role—exert only marginal influence. These findings indicate fundamental deficiencies in current LLMs’ syntactic representations. The work introduces RC attachment ambiguity as a novel cross-lingual benchmark for evaluating linguistic knowledge and cognitive preferences in LLMs, providing critical empirical evidence on the boundaries of their syntactic competence.

Technology Category

Application Category

📝 Abstract
This paper leverages past sentence processing studies to investigate whether monolingual and multilingual LLMs show human-like preferences when presented with examples of relative clause attachment ambiguities in Italian and English. Furthermore, we test whether these preferences can be modulated by lexical factors (the type of verb/noun in the matrix clause) which have been shown to be tied to subtle constraints on syntactic and semantic relations. Our results overall showcase how LLM behavior varies interestingly across models, but also general failings of these models in correctly capturing human-like preferences. In light of these results, we argue that RC attachment is the ideal benchmark for cross-linguistic investigations of LLMs' linguistic knowledge and biases.
Problem

Research questions and friction points this paper is trying to address.

Investigates syntactic biases in multilingual transformers using RC ambiguities.
Tests if LLMs show human-like preferences in RC attachment ambiguities.
Evaluates lexical factors' impact on syntactic and semantic relations in LLMs.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Investigates multilingual LLMs with RC ambiguities
Tests lexical factors' impact on syntactic preferences
Proposes RC attachment as LLM benchmark