🤖 AI Summary
This study empirically tests the classical linguistic hypothesis—originating in generative grammar (Chomsky & Pinker, 1994)—that “the vast majority of sentences are novel combinations.” Using a multi-genre corpus spanning news, dialogue, academic writing, and social media, we conduct exact string-level sentence matching and frequency counting via NLTK. Results reveal that while novel sentences predominate (50%–90%), repeated sentences occur significantly across all genres (6.2%–48.7%), with repetition rates strongly genre-dependent: highest in dialogue and social media, lowest in news and academic texts. This is the first systematic cross-genre empirical investigation of sentence uniqueness. It demonstrates that sentence novelty is not an absolute linguistic universal but is profoundly constrained by pragmatic function and genre-specific conventions—challenging the formal-linguistic emphasis on unbounded syntactic productivity. The findings underscore the necessity of integrating usage-based and genre-aware perspectives into theories of sentence production and linguistic competence.
📝 Abstract
A repeated claim in linguistics is that the majority of linguistic utterances are unique. For example, Pinker (1994: 10), summarizing an argument by Noam Chomsky, states that "virtually every sentence that a person utters or understands is a brand-new combination of words, appearing for the first time in the history of the universe." With the increased availability of large corpora, this is a claim that can be empirically investigated. The current paper addresses the question by using the NLTK Python library to parse corpora of different genres, providing counts of exact string matches in each. Results show that while completely unique sentences are often the majority of corpora, this is highly constrained by genre, and that duplicate sentences are not an insignificant part of any individual corpus.