Clarifying orthography: Orthographic transparency as compressibility

📅 2025-05-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Orthographic transparency lacks cross-script comparable quantitative metrics. Method: We propose a novel measure grounded in algorithmic information theory, defining transparency as mutual compressibility between orthographic and phonemic strings. This unsupervised approach models transparency as mutual compression length—unifying the quantification of both spelling irregularities and rule complexity—and applies universally across alphabetic, syllabic, logographic, and other writing systems. Technically, we integrate pretrained code-length estimators with neural language models (e.g., Transformers) to end-to-end estimate compression distances from large-scale text–speech pairs. Results: Evaluated on 22 languages spanning five major script types, our metric aligns with linguistic intuition and demonstrates strong interpretability, computational tractability, and cross-script comparability—establishing the first theoretically principled, general-purpose, and reproducible quantitative benchmark for cross-linguistic orthographic analysis.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Evaluation and AnalysisSearch and Optimization: Evaluation and Analysis

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsResponsible Web: Algorithmic accountability and transparency on the web
📝 Abstract
Orthographic transparency -- how directly spelling is related to sound -- lacks a unified, script-agnostic metric. Using ideas from algorithmic information theory, we quantify orthographic transparency in terms of the mutual compressibility between orthographic and phonological strings. Our measure provides a principled way to combine two factors that decrease orthographic transparency, capturing both irregular spellings and rule complexity in one quantity. We estimate our transparency measure using prequential code-lengths derived from neural sequence models. Evaluating 22 languages across a broad range of script types (alphabetic, abjad, abugida, syllabic, logographic) confirms common intuitions about relative transparency of scripts. Mutual compressibility offers a simple, principled, and general yardstick for orthographic transparency.
Problem

Research questions and friction points this paper is trying to address.

Lack of unified metric for orthographic transparency across scripts
Quantifying transparency via orthographic-phonological compressibility
Measuring irregular spellings and rule complexity in one quantity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantify transparency via mutual compressibility
Combine irregular spellings and rule complexity
Estimate using neural sequence models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.