arabic text preprocessing

Designs and implements preprocessing pipelines and software components that clean, normalize, tokenize, and segment Arabic-language text for downstream NLP models and analyses. Work includes handling Arabic-script normalization (orthographic variants, diacritics, hamza/alef forms), clitic and morphological segmentation, stopword removal, encoding/Unicode issues, and tokenization tuned for statistical or neural systems.

arabictextpreprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3

Dec 20, 2025
MK
Mark Kashirskiy
🏛️ Higher School of Economics | Saint Petersburg State University

To address the high redundancy and low compression ratio of generic tokenizers on morphologically rich Arabic, this paper proposes an Arabic-optimized SentencePiece Unigram tokenizer and a Language Expansion Pipeline (LEP). We design a dedicated Arabic normalization pipeline—unifying Alif variants, removing diacritics, and normalizing Arabic–Indic numerals—and introduce LEP: initializing new tokens via mean subword embeddings and enabling efficient vocabulary expansion through selective Transformer layer unfreezing. Integrated into Qwen3-0.6B, our tokenizer reduces tokenization fertility by 18% (from 1.199 to 1.35 tokens/word) and lowers validation loss from 8.28 to 2.43 after 800 training steps on 100K samples. All code, the optimized tokenizer, and model checkpoints are publicly released.

Addresses orthographic variations like Alif and diacriticsIntegrates tokenizer into Qwen3 via vocabulary extensionOptimizes Arabic tokenization for LLMs with normalization

This study addresses the lack of systematic investigation into the ordering of preprocessing steps in sentiment analysis, which has constrained model performance and efficiency. Focusing on Twitter data, it presents the first quantitative evaluation of the relative impact and optimal sequencing of key preprocessing techniques—including tokenization, text cleaning, stemming, stopword removal (with negations preserved), and spelling correction. Through comprehensive combinatorial experiments, the work demonstrates that tokenization contributes most significantly to model performance, while spelling correction has the least effect. The identified optimal pipeline—tokenization followed by text cleaning, stemming, and stopword removal—substantially enhances model effectiveness and reduces trial-and-error costs, establishing a reproducible and efficient preprocessing paradigm for sentiment analysis.

machine learningpreprocessingsentiment analysis

Investigating Large Language Models' Linguistic Abilities for Text Preprocessing

Oct 13, 2025
MB
Marco Braga
🏛️ University of Milano-Bicocca | Politecnico di Torino

Conventional text preprocessing techniques—such as stopword removal, lemmatization, and stemming—rely heavily on language-specific linguistic rules and ignore contextual information, limiting their generalizability across multilingual settings. Method: This paper pioneers a systematic investigation of large language models (LLMs) as context-aware, universal preprocessors. Leveraging prompt engineering, we uniformly perform the three preprocessing tasks across six European languages without language-specific annotations or handcrafted rules. Contribution/Results: Experiments show LLMs achieve 97%, 82%, and 74% accuracy on stopword removal, lemmatization, and stemming, respectively. Downstream text classification models fed with LLM-preprocessed inputs attain up to a 6-percentage-point improvement in F1 score. This work demonstrates the feasibility and effectiveness of LLM-driven, end-to-end, context-sensitive, and multilingual-compatible text preprocessing—establishing a novel paradigm that reduces reliance on manual linguistic rules and enhances preprocessing robustness.

LLMs address context-dependent text preprocessing limitationsLLMs improve text classification accuracy over traditional techniquesTraditional methods ignore contextual information in preprocessing

Tahakom LLM guidelines and receipts: from pre-training data to an Arabic LLM

Oct 15, 2025
AA
Areej AlOtaibi
🏛️ University of Oxford

This work addresses three core challenges in developing Arabic large language models (LLMs): high data noise, inadequate tokenizer adaptation, and weak evaluation frameworks. To tackle these, we propose a systematic solution comprising: (1) a multi-stage Arabic-specific data cleaning framework; (2) a hybrid tokenizer integrating morphology-aware segmentation with subword tokenization; and (3) a multidimensional evaluation benchmark covering linguistic understanding, generation quality, and cultural appropriateness. Leveraging a massive, high-quality Arabic corpus, we implement an end-to-end customized pretraining pipeline. The resulting open-source foundational model achieves an average 12.3% improvement over state-of-the-art open Arabic LLMs across major Arabic benchmarks, with substantially enhanced reasoning and generation capabilities. We publicly release the cleaned dataset, tokenizer toolkit, and evaluation suite to foster sustainable advancement of the Arabic AI ecosystem.

Addressing data curation challenges for Arabic language modelsEvaluating tokenizer design impact on Arabic model performanceProposing systematic corrections for Arabic evaluation frameworks

Large Language Models and Arabic Content: A Review

May 12, 2025
HR
Haneh Rhel
🏛️ University of Strathclyde

Arabic NLP has long suffered from scarce resources, dialectal diversity, rich morphology, and pervasive orthographic variation. This paper presents the first systematic survey of large language models (LLMs) for Arabic processing, covering Arabic-specific pretraining, multidiaglect adaptation strategies, supervised/instruction fine-tuning, and prompt engineering techniques. It integrates major evaluation benchmarks—including ArabicMMLU and AQAD—to assess model capabilities across linguistic dimensions. The study elucidates how multilingual pretraining enhances morphological generalization and orthographic variant modeling in Arabic, identifies critical bottlenecks (e.g., inadequate low-resource dialect coverage and evaluation dataset bias), and pinpoints essential data gaps. Our analysis yields a technical roadmap for Arabic AI resource development, supporting the advancement of robust, multi-standard-compliant Arabic NLP systems.

Addressing scarcity of Arabic resources for LLMsEnhancing LLM performance for Arabic dialects and tasksOvercoming Arabic NLP challenges like complex morphology

Latest Papers

What's happening recently
View more

This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root-pattern morphology, probing whether they capture genuine morphological structure or rely on surface memorization. Arabic morphological system provides a rich testbed for analyzing how LLMs handle complex, non-concatenative forms and how tokenization choices influence this process. Our study begins with an evaluation of morphological fidelity across Arabic and multilingual tokenizers against gold-standard segmentation, followed by an analysis of LLM performance in productive root-pattern generation using a newly developed test set. Our findings across seven Arabic-centric and multilingual LLMs and their respective tokenizers reveal that tokenizer morphological alignment is not necessary nor sufficient for morphological generation, which questions the role of morphological tokenization in downstream performance.

This study addresses the underperformance of large language models (LLMs) on morphosyntactic tagging and dependency parsing tasks in Modern Standard Arabic, a language characterized by complex morphology and orthographic ambiguity. It presents the first systematic evaluation of instruction-tuned LLMs under zero-shot prompting and retrieval-augmented in-context learning (ICL) settings. Through carefully designed prompts and example selection strategies, experiments are conducted on the Arabic Treebank. Results show that LLMs achieve near-supervised performance on morphological feature tagging and match specialized parsers in dependency parsing. Notably, retrieval-based ICL substantially improves tokenization, tagging, and parsing accuracy on raw text, highlighting both the potential and limitations of LLMs in handling intricate morphosyntactic interactions.

Arabicdependency parsinglarge language models

Existing Arabic–English dictionaries are predominantly available in unstructured textual formats, rendering them unsuitable for direct use in natural language processing tasks. This work addresses the absence of standardized microstructural organization in the widely used *Al-Mawrid* dictionary by proposing a semi-automatic structuring approach grounded in inductive rules, achieving the first hierarchical parsing of its lexical entries. The method employs Parsing Expression Grammars (PEG) to construct a core parser that explicitly models subentries, definition phrases, domain labels, cross-references, and translation equivalents. Experimental results demonstrate that the proposed pipeline effectively converts the dictionary into a machine-readable format, exhibiting both practical feasibility and reasonable accuracy in structuring non-standardized lexical resources.

Arabic-English dictionarydictionary structuringlexical information

Hot Scholars

MA

Mohammad AL-Smadi

Associate Professor, Qatar University
Natural Language ProcessingArtificial IntelligenceTechnology-enhanced learningSemantic Computing
BS

Bilal Sowan

University of Petra
Data MiningMachine LearningData ScienceBusiness Intelligence
WB

Wadii Boulila

Professor of Computer Science, Leader of Robotics & Internet of Things Lab, Prince Sultan University
Data ScienceMachine LearningUncertainty ModelingRemote Sensing
SA

Saleh Almohaimeed

King Saud University
Deep LearningNatural Language ProcessingSemantic Parsing