🤖 AI Summary
This study investigates the regularities and underlying causes of phoneme frequency distributions across the world’s languages. By developing an information-theoretic framework that integrates macro- and micro-level perspectives, the authors propose modeling cross-linguistic phoneme frequencies with a symmetric Dirichlet distribution and incorporate articulatory, phonological, and lexical constraints through a maximum entropy model to predict language-specific phoneme probabilities. The research reveals a compensatory relationship between phonemic inventory size and relative entropy. The proposed model not only accurately captures the rank–frequency distribution of phonemes but also effectively predicts language-specific phoneme usage patterns, thereby demonstrating the explanatory power of information theory in accounting for both phonological universals and diversity.
📝 Abstract
We demonstrate that the frequency distribution of phonemes across languages can be explained at both macroscopic and microscopic levels. Macroscopically, phoneme rank-frequency distributions closely follow the order statistics of a symmetric Dirichlet distribution whose single concentration parameter scales systematically with phonemic inventory size, revealing a robust compensation effect whereby larger inventories exhibit lower relative entropy. Microscopically, a Maximum Entropy model incorporating constraints from articulatory, phonotactic, and lexical structure accurately predicts language-specific phoneme probabilities. Together, these findings provide a unified information-theoretic account of phoneme frequency structure.