LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the "stochastic parrot" hypothesis by investigating whether large language models (LLMs) can transcend statistical pattern matching to achieve semantically mediated abstract reasoning. We propose a novel testing paradigm based on constructed languages featuring counterintuitive feature combinations, evaluating the models' zero-shot inference capabilities through natural language rule descriptions without exemplars. Experimental results demonstrate that LLMs exhibit consistent rule-following behavior, confirming that statistical learning can drive the emergence of semantic abstraction. This work provides compelling evidence against the strong stochastic parrot hypothesis, offering new empirical insights and perspectives for elucidating the underlying mechanisms of abstract reasoning in large language models.
📝 Abstract
The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.
Problem

Research questions and friction points this paper is trying to address.

stochastic parrot hypothesis
large language models
abstraction
statistical pattern matching
meaning-mediated abstraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Stochastic Parrot Hypothesis
Conlang-like Tasks
Meaning-mediated Abstraction
Statistical Pattern Matching
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Julia Witte Zimmerman
Computational Story Lab, Vermont Complex Systems Institute, University of Vermont, Burlington, VT 05405, USA
C
Calla G. Beauregard
Computational Story Lab, Vermont Complex Systems Institute, University of Vermont, Burlington, VT 05405, USA
Tabia Tanzin Prama
Tabia Tanzin Prama
Phd Student of Computer Science
Data MiningNLPHealth InformaticsAI Ethics
P
Parisa Suchdev
Computational Ethics Lab, Vermont Complex Systems Institute, University of Vermont, Burlington, VT 05405, USA
K
Kathryn Cramer
Computational Story Lab, Vermont Complex Systems Institute, University of Vermont, Burlington, VT 05405, USA
E
Elisabeth Kollrack
Computational Story Lab, Vermont Complex Systems Institute, University of Vermont, Burlington, VT 05405, USA