🤖 AI Summary
This study investigates large language models’ (LLMs) capacity to model and generate iconicity—the direct perceptual correspondence between word form and meaning. Methodologically, we employed GPT-4 to generate iconic pseudowords, validated them through cross-linguistic human experiments (Czech, German), and conducted introspective LLM evaluations (GPT-4, Claude 3.5 Sonnet). Our results demonstrate, for the first time, that LLMs not only implicitly encode iconic patterns present in natural language but also actively construct artificial lexicons exhibiting higher iconic strength than natural languages. Moreover, LLMs themselves achieve significantly higher accuracy than humans in meaning inference tasks involving these pseudowords; correspondingly, human participants guess the meanings of such LLM-generated pseudowords with markedly greater accuracy than they do for unfamiliar natural-language words. These findings reveal that LLMs possess superior intrinsic iconic modeling and generalization capabilities compared to humans—providing novel evidence for theories of language evolution, human–AI interaction, and AI symbol grounding.
📝 Abstract
Lexical iconicity, a direct relation between a word's meaning and its form, is an important aspect of every natural language, most commonly manifesting through sound-meaning associations. Since Large language models' (LLMs') access to both meaning and sound of text is only mediated (meaning through textual context, sound through written representation, further complicated by tokenization), we might expect that the encoding of iconicity in LLMs would be either insufficient or significantly different from human processing. This study addresses this hypothesis by having GPT-4 generate highly iconic pseudowords in artificial languages. To verify that these words actually carry iconicity, we had their meanings guessed by Czech and German participants (n=672) and subsequently by LLM-based participants (generated by GPT-4 and Claude 3.5 Sonnet). The results revealed that humans can guess the meanings of pseudowords in the generated iconic language more accurately than words in distant natural languages and that LLM-based participants are even more successful than humans in this task. This core finding is accompanied by several additional analyses concerning the universality of the generated language and the cues that both human and LLM-based participants utilize.