Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the relationship between Arabic script forms and their linguistic functions relies on visual iconicity or can be effectively replaced by arbitrary yet consistent character mappings. To this end, we randomly remap Arabic letters onto 19 dotless base glyphs (rasm) and systematically compare the performance of original Arabic script, standard dotless forms, and fully randomized mappings across multiple NLP tasks—including language modeling, text classification, sequence labeling, machine translation, and original text recovery—under both word-level and character-level tokenization strategies. From 2,000 random mappings, high- and low-entropy representatives are selected for evaluation. Results demonstrate that randomized mappings achieve performance comparable to standard Arabic while substantially reducing vocabulary size, out-of-vocabulary rates, model scale, and training costs, thereby revealing the fundamentally arbitrary nature of the form–function correspondence in Arabic orthography.
📝 Abstract
Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscripts were written without dots yet remained interpretable, dot removal offers a natural test of whether these visual distinctions are functionally necessary. Prior work has shown that dotless Arabic can remain readable and effective for natural language processing (NLP), but it remains unclear whether this success depends on preserving the original rasm groupings or whether arbitrary but consistent remappings to the same reduced rasm set can achieve comparable performance. We address this question by comparing standard dotted and dotless Arabic with arbitrary character remappings constrained to the same 19 undotted rasms. We generated 2,000 random remappings under word- and character-level tokenization and selected four representative mappings with the highest and lowest entropy values. These representations were evaluated across language modeling, text classification, sequence labeling, machine translation, and restoration to the original script. The results show that neither preserving original character distinctions nor retaining traditional rasm-based groupings is necessary for strong NLP performance. Random remappings achieve competitive performance while reducing vocabulary size, out-of-vocabulary (OOV) rates, model size, and training cost. These findings suggest that, from an NLP perspective, Arabic character form-function relationships are largely arbitrary: models rely more on stable distributional structure than on the visual iconicity of letter forms.
Problem

Research questions and friction points this paper is trying to address.

Arabic NLP
character iconicity
arbitrariness
rasm
dotless Arabic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Arabic NLP
character remapping
rasm
iconicity vs. arbitrariness
dotless Arabic
🔎 Similar Papers
No similar papers found.