Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models (LLMs) exhibit human-like sublexical sensitivity when processing pseudowords. Employing a two-alternative forced-choice paradigm combined with cosine similarity analysis and inference token consumption monitoring, the authors systematically compare the performance of LLMs, human participants, and fastText on an Italian pseudoword task. Results indicate that LLMs underperform fastText in pure pseudoword conditions and fail to replicate the sublexical cue alignment patterns characteristic of human processing. These findings reveal that LLMs lack human-like sublexical processing mechanisms, identifying tokenization strategies and training data coverage as primary constraints. Overall, this work provides novel evidence for understanding the morphological processing limitations inherent in current LLM architectures.
📝 Abstract
Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseudoword processing. Yet whether LLMs exhibit comparable sensitivity to these cues remains unclear. We tested five LLMs on two Italian two-alternative forced-choice pseudoword experiments and compared their responses with a human behavioural baseline. LLMs aligned more reliably with humans when real-word options provided a lexical familiarity cue than in the pseudoword-only condition, where they fell substantially below fastText, a character-n-gram model. In addition, the sublexical cosine-similarity cue that reliably drove human--fastText agreement did not consistently transfer to human--LLM alignment, and reasoning-token expenditure bore no consistent relation to human processing difficulty. These findings suggest that LLMs do not necessarily share the sublexical cues that govern human pseudoword processing; we discuss tokenization and training-data coverage as candidate explanations.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
pseudoword processing
sublexical sensitivity
systematicity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pseudoword processing
Sublexical sensitivity
Large Language Models
Tokenization
Character n-gram model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jing Chen
Department of Psychology, University of Milano-Bicocca, Piazza dell’Ateneo Nuovo 1, 20126 Milano, Italy
G
Giulia Loca
Department of Psychology, University of Milano-Bicocca, Piazza dell’Ateneo Nuovo 1, 20126 Milano, Italy
S
Simona Amenta
Department of Informatics, Systems and Communication – DISCo, University of Milano-Bicocca, Milano, Italy
Marco Marelli
Marco Marelli
Department of Psychology, University of Milano-Bicocca