text processing

Designs, builds, and evaluates pipelines and software components that ingest, clean, normalize, segment, tokenize, annotate, and transform raw text into structured representations for downstream use. This includes implementing and analyzing algorithms and tools for tokenization, morphological processing (stemming/lemmatization), sentence splitting, parsing, tagging (e.g., POS/NER), indexing, and extraction of features or information from textual data.

textprocessing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Investigating Large Language Models' Linguistic Abilities for Text Preprocessing

Oct 13, 2025
MB
Marco Braga
🏛️ University of Milano-Bicocca | Politecnico di Torino

Conventional text preprocessing techniques—such as stopword removal, lemmatization, and stemming—rely heavily on language-specific linguistic rules and ignore contextual information, limiting their generalizability across multilingual settings. Method: This paper pioneers a systematic investigation of large language models (LLMs) as context-aware, universal preprocessors. Leveraging prompt engineering, we uniformly perform the three preprocessing tasks across six European languages without language-specific annotations or handcrafted rules. Contribution/Results: Experiments show LLMs achieve 97%, 82%, and 74% accuracy on stopword removal, lemmatization, and stemming, respectively. Downstream text classification models fed with LLM-preprocessed inputs attain up to a 6-percentage-point improvement in F1 score. This work demonstrates the feasibility and effectiveness of LLM-driven, end-to-end, context-sensitive, and multilingual-compatible text preprocessing—establishing a novel paradigm that reduces reliance on manual linguistic rules and enhances preprocessing robustness.

LLMs address context-dependent text preprocessing limitationsLLMs improve text classification accuracy over traditional techniquesTraditional methods ignore contextual information in preprocessing

Splintering Nonconcatenative Languages for Better Tokenization

Mar 18, 2025
BG
Bar Gazit
🏛️ Ben-Gurion University of the Negev | DICTA

Conventional subword tokenization algorithms (e.g., BPE, UnigramLM) assume linear concatenative morphology, making them ill-suited for non-concatenative languages—such as Hebrew and Arabic (root-and-pattern morphology) or agglutinative yet discontinuous languages like Malay—where morphological units are not orthographically contiguous. This leads to poor preservation of morphological integrity during tokenization. Method: We propose SPLINTER, the first preprocessing framework that systematically models non-concatenative morphology’s impact on tokenization. It performs language-aware text linearization by integrating morphological rules with statistical heuristics to reorder morphemic units into contiguous sequences amenable to standard subword algorithms. Contribution/Results: SPLINTER is compatible with both BPE and UnigramLM. Evaluated within the multilingual BERT framework, it improves morphological integrity by +23% on Hebrew, Arabic, and Malay, and yields an average +1.8 percentage point gain on Hebrew downstream tasks—effectively breaking the implicit concatenative-morphology assumption underlying mainstream tokenizers.

Addresses tokenization challenges in nonconcatenative languages.Improves tokenizer performance in Hebrew, Arabic, and Malay.Proposes SPLINTER for better representation of morphological patterns.

This study addresses the lack of systematic investigation into the ordering of preprocessing steps in sentiment analysis, which has constrained model performance and efficiency. Focusing on Twitter data, it presents the first quantitative evaluation of the relative impact and optimal sequencing of key preprocessing techniques—including tokenization, text cleaning, stemming, stopword removal (with negations preserved), and spelling correction. Through comprehensive combinatorial experiments, the work demonstrates that tokenization contributes most significantly to model performance, while spelling correction has the least effect. The identified optimal pipeline—tokenization followed by text cleaning, stemming, and stopword removal—substantially enhances model effectiveness and reduces trial-and-error costs, establishing a reproducible and efficient preprocessing paradigm for sentiment analysis.

machine learningpreprocessingsentiment analysis

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Oct 28, 2024
QZ
Qintong Zhang
🏛️ Shanghai Artificial Intelligence Laboratory | Peking University

This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.

Address challenges in layout detection and multi-modal data integrationConvert unstructured documents into structured machine-readable dataImprove parsing accuracy for complex layouts and high-density text

Latest Papers

What's happening recently
View more

Tokens with Meaning: A Hybrid Tokenization Approach for NLP

Aug 19, 2025
MA
M. Ali Bayram
🏛️ Yıldız Technical University | Yeditepe University | University of Chicago | Istanbul Bilgi University

Existing subword tokenization methods (e.g., BPE, WordPiece) rely heavily on surface-form frequency statistics, rendering them ill-suited for morphologically rich, agglutinative languages. To address this, we propose a hybrid tokenization framework integrating linguistic rules with statistical learning: (1) phonemic normalization mitigates orthographic variation; (2) an explicit root–affix dictionary models morphological structure; (3) a shared identifier mechanism balances morpheme fidelity and subword efficiency; and (4) unified special tokens handle whitespace and case, curbing vocabulary bloat. Evaluated on the TR-MMLU benchmark, our method achieves 90.29% tokenization accuracy and 85.8% pure segmentation rate for Turkish—substantially outperforming LLaMA, Gemma, and GPT tokenizers. This work represents the first systematic integration of phonological, morphological, and statistical modeling in large-scale pretraining tokenization, significantly enhancing cross-lingual semantic consistency and out-of-vocabulary generalization—particularly for agglutinative languages.

Combining linguistic rules with statistical subword segmentationImproving tokenization for morphologically rich languagesReducing vocabulary redundancy while preserving semantic meaning

Tokenization strategies significantly impact the modeling of assembly code, yet their influence on downstream tasks—such as function signature prediction—and trade-offs among intrinsic properties (e.g., vocabulary coverage, semantic fidelity) remain underexplored. Method: We systematically evaluate byte-level, subword-level, and custom tokenizers—including Byte-Pair Encoding (BPE), WordPiece, and assembly-aware variants—using Llama-3.2, BERT, and BART within a unified framework. Our analysis integrates vocabulary compression profiling, representation fidelity measurement, and assembly-specific preprocessing rules. Contribution/Results: Tokenizer choice critically affects predictive performance; certain intrinsic metrics (e.g., opcode coverage, mnemonic preservation) correlate strongly with downstream accuracy. We propose a lightweight, binary-code-oriented tokenization optimization pathway that enhances semantic capture without compromising computational efficiency. This work establishes a reproducible evaluation paradigm and practical tokenizer design guidelines for adapting large language models to low-level code.

Analyzing trade-offs between intrinsic tokenizer properties and practical utilityAssessing tokenizer effects on downstream tasks like function signature predictionEvaluating how tokenization algorithms impact assembly code analysis performance

Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi

Aug 18, 2025
MA
M. Ali Bayram
🏛️ Yıldız Technical University | Yeditepe University | The University of Chicago | Istanbul Bilgi University

This study addresses the tokenization challenge for morphologically rich, low-resource languages (e.g., Turkish). We propose the first multidimensional evaluation framework tailored to such languages, built upon the newly constructed TR-MMLU dataset. We systematically benchmark mainstream tokenizers across four dimensions: vocabulary size, token count, processing efficiency, and preservation of linguistic structure—introducing two novel metrics: language-specific tokenization ratio (%TR) and token purity (%Pure). Results show that %TR strongly correlates with downstream task performance, whereas %Pure exhibits weak correlation; scaling model parameters alone does not improve linguistic understanding. Crucially, this work provides the first empirical validation that %TR is a more effective indicator of tokenization quality than conventional metrics. It underscores the necessity of linguistically grounded, language-specific tokenization strategies for morphologically complex languages and establishes a reproducible evaluation paradigm and methodological foundation for low-resource NLP.

Assessing correlation between tokenization and model performanceEvaluating tokenization impact on Turkish language modelsProposing metrics for linguistic structure preservation

A Task-Oriented Evaluation Framework for Text Normalization in Modern NLP Pipelines

Nov 25, 2025
MA
Md Abdullah Al Kafi
🏛️ Daffodil International University | Alliance University

Existing stemmer evaluation methods fail to quantify the semantic degradation caused by over-stemming in downstream tasks. This paper proposes the first task-oriented text normalization evaluation framework, overcoming the limitations of traditional lemmatization-based metrics by jointly measuring three dimensions: Stemming Effectiveness Score (SES), Model Performance Delta (MPD), and Average Normalized Levenshtein Distance (ANLD). The framework enables integrated analysis of both efficiency and semantic safety—the first of its kind. Empirical evaluation across multiple languages reveals a critical insight: high stemming recall does not necessarily improve downstream performance. For instance, the Bangla stemmer suffers performance degradation due to aggressive over-stemming, whereas the English Snowball stemmer achieves a superior trade-off between effectiveness and semantic fidelity. This work establishes a principled, task-aware methodology for evaluating and comparing stemmers beyond surface-form matching.

Assessing stemming impact on downstream task performanceEvaluating stemming methods' utility and safety in NLP pipelinesMeasuring semantic preservation during word normalization processes

This work identifies and systematically characterizes the phenomenon of “intermediate merge residues” in Byte Pair Encoding (BPE) tokenizers—low-frequency subword tokens generated during training that are rarely used during inference. These redundant tokens unnecessarily consume vocabulary capacity and degrade model robustness to input noise and spelling errors. To address this issue, the authors propose LiteToken, a lightweight post-training optimization method that directly prunes existing BPE vocabularies without requiring retraining. By analyzing token formation mechanisms and empirical usage frequencies, LiteToken effectively removes superfluous tokens, substantially reducing vocabulary fragmentation and model parameter count while preserving original task performance and enhancing robustness to anomalous inputs.

BPE tokenizationintermediate merge residuestoken fragmentation

Hot Scholars

PN

Ping Nie

Waterloo University
Natural Language ProcessingInformation RetrievalRecommendation SystemsTime Series Forecasting
CW

Cong Wei

University of Waterloo
ReasoningDiffusionEfficiency
JF

Julian Frattini

University of Gothenburg | Chalmers University of Technology
Hybrid AI Software SystemsRequirements EngineeringResearch Methodology
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
IB

Iris Beerepoot

Utrecht University
workaroundsprocess miningbusiness process managementhealth care