🤖 AI Summary
Orthographic transparency lacks cross-script comparable quantitative metrics. Method: We propose a novel measure grounded in algorithmic information theory, defining transparency as mutual compressibility between orthographic and phonemic strings. This unsupervised approach models transparency as mutual compression length—unifying the quantification of both spelling irregularities and rule complexity—and applies universally across alphabetic, syllabic, logographic, and other writing systems. Technically, we integrate pretrained code-length estimators with neural language models (e.g., Transformers) to end-to-end estimate compression distances from large-scale text–speech pairs. Results: Evaluated on 22 languages spanning five major script types, our metric aligns with linguistic intuition and demonstrates strong interpretability, computational tractability, and cross-script comparability—establishing the first theoretically principled, general-purpose, and reproducible quantitative benchmark for cross-linguistic orthographic analysis.
📝 Abstract
Orthographic transparency -- how directly spelling is related to sound -- lacks a unified, script-agnostic metric. Using ideas from algorithmic information theory, we quantify orthographic transparency in terms of the mutual compressibility between orthographic and phonological strings. Our measure provides a principled way to combine two factors that decrease orthographic transparency, capturing both irregular spellings and rule complexity in one quantity. We estimate our transparency measure using prequential code-lengths derived from neural sequence models. Evaluating 22 languages across a broad range of script types (alphabetic, abjad, abugida, syllabic, logographic) confirms common intuitions about relative transparency of scripts. Mutual compressibility offers a simple, principled, and general yardstick for orthographic transparency.