Are most sentences unique? An empirical examination of Chomskyan claims

📅 2025-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study empirically tests the classical linguistic hypothesis—originating in generative grammar (Chomsky & Pinker, 1994)—that “the vast majority of sentences are novel combinations.” Using a multi-genre corpus spanning news, dialogue, academic writing, and social media, we conduct exact string-level sentence matching and frequency counting via NLTK. Results reveal that while novel sentences predominate (50%–90%), repeated sentences occur significantly across all genres (6.2%–48.7%), with repetition rates strongly genre-dependent: highest in dialogue and social media, lowest in news and academic texts. This is the first systematic cross-genre empirical investigation of sentence uniqueness. It demonstrates that sentence novelty is not an absolute linguistic universal but is profoundly constrained by pragmatic function and genre-specific conventions—challenging the formal-linguistic emphasis on unbounded syntactic productivity. The findings underscore the necessity of integrating usage-based and genre-aware perspectives into theories of sentence production and linguistic competence.

Technology Category

Natural Language Processing: GenerationCognitive Modeling & Cognitive Systems: Computational CreativityConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Web Mining and Content Analysis: Robustness and generalizability of Web mining methodsSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSecurity and Privacy: Large-scale security measurements
📝 Abstract
A repeated claim in linguistics is that the majority of linguistic utterances are unique. For example, Pinker (1994: 10), summarizing an argument by Noam Chomsky, states that "virtually every sentence that a person utters or understands is a brand-new combination of words, appearing for the first time in the history of the universe." With the increased availability of large corpora, this is a claim that can be empirically investigated. The current paper addresses the question by using the NLTK Python library to parse corpora of different genres, providing counts of exact string matches in each. Results show that while completely unique sentences are often the majority of corpora, this is highly constrained by genre, and that duplicate sentences are not an insignificant part of any individual corpus.
Problem

Research questions and friction points this paper is trying to address.

Empirically tests Chomsky's claim about sentence uniqueness
Analyzes corpus data using NLTK to find duplicate sentences
Examines how sentence uniqueness varies across different genres
Innovation

Methods, ideas, or system contributions that make the work stand out.

Used NLTK Python library for corpus parsing
Counted exact string matches across genres
Analyzed sentence uniqueness frequency variations
🔎 Similar Papers
No similar papers found.