Score
Designs, builds, and evaluates pipelines and software components that ingest, clean, normalize, segment, tokenize, annotate, and transform raw text into structured representations for downstream use. This includes implementing and analyzing algorithms and tools for tokenization, morphological processing (stemming/lemmatization), sentence splitting, parsing, tagging (e.g., POS/NER), indexing, and extraction of features or information from textual data.
Conventional text preprocessing techniques—such as stopword removal, lemmatization, and stemming—rely heavily on language-specific linguistic rules and ignore contextual information, limiting their generalizability across multilingual settings. Method: This paper pioneers a systematic investigation of large language models (LLMs) as context-aware, universal preprocessors. Leveraging prompt engineering, we uniformly perform the three preprocessing tasks across six European languages without language-specific annotations or handcrafted rules. Contribution/Results: Experiments show LLMs achieve 97%, 82%, and 74% accuracy on stopword removal, lemmatization, and stemming, respectively. Downstream text classification models fed with LLM-preprocessed inputs attain up to a 6-percentage-point improvement in F1 score. This work demonstrates the feasibility and effectiveness of LLM-driven, end-to-end, context-sensitive, and multilingual-compatible text preprocessing—establishing a novel paradigm that reduces reliance on manual linguistic rules and enhances preprocessing robustness.
Conventional subword tokenization algorithms (e.g., BPE, UnigramLM) assume linear concatenative morphology, making them ill-suited for non-concatenative languages—such as Hebrew and Arabic (root-and-pattern morphology) or agglutinative yet discontinuous languages like Malay—where morphological units are not orthographically contiguous. This leads to poor preservation of morphological integrity during tokenization. Method: We propose SPLINTER, the first preprocessing framework that systematically models non-concatenative morphology’s impact on tokenization. It performs language-aware text linearization by integrating morphological rules with statistical heuristics to reorder morphemic units into contiguous sequences amenable to standard subword algorithms. Contribution/Results: SPLINTER is compatible with both BPE and UnigramLM. Evaluated within the multilingual BERT framework, it improves morphological integrity by +23% on Hebrew, Arabic, and Malay, and yields an average +1.8 percentage point gain on Hebrew downstream tasks—effectively breaking the implicit concatenative-morphology assumption underlying mainstream tokenizers.
This study addresses the lack of systematic investigation into the ordering of preprocessing steps in sentiment analysis, which has constrained model performance and efficiency. Focusing on Twitter data, it presents the first quantitative evaluation of the relative impact and optimal sequencing of key preprocessing techniques—including tokenization, text cleaning, stemming, stopword removal (with negations preserved), and spelling correction. Through comprehensive combinatorial experiments, the work demonstrates that tokenization contributes most significantly to model performance, while spelling correction has the least effect. The identified optimal pipeline—tokenization followed by text cleaning, stemming, and stopword removal—substantially enhances model effectiveness and reduces trial-and-error costs, establishing a reproducible and efficient preprocessing paradigm for sentiment analysis.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
Existing subword tokenization methods (e.g., BPE, WordPiece) rely heavily on surface-form frequency statistics, rendering them ill-suited for morphologically rich, agglutinative languages. To address this, we propose a hybrid tokenization framework integrating linguistic rules with statistical learning: (1) phonemic normalization mitigates orthographic variation; (2) an explicit root–affix dictionary models morphological structure; (3) a shared identifier mechanism balances morpheme fidelity and subword efficiency; and (4) unified special tokens handle whitespace and case, curbing vocabulary bloat. Evaluated on the TR-MMLU benchmark, our method achieves 90.29% tokenization accuracy and 85.8% pure segmentation rate for Turkish—substantially outperforming LLaMA, Gemma, and GPT tokenizers. This work represents the first systematic integration of phonological, morphological, and statistical modeling in large-scale pretraining tokenization, significantly enhancing cross-lingual semantic consistency and out-of-vocabulary generalization—particularly for agglutinative languages.
Tokenization strategies significantly impact the modeling of assembly code, yet their influence on downstream tasks—such as function signature prediction—and trade-offs among intrinsic properties (e.g., vocabulary coverage, semantic fidelity) remain underexplored. Method: We systematically evaluate byte-level, subword-level, and custom tokenizers—including Byte-Pair Encoding (BPE), WordPiece, and assembly-aware variants—using Llama-3.2, BERT, and BART within a unified framework. Our analysis integrates vocabulary compression profiling, representation fidelity measurement, and assembly-specific preprocessing rules. Contribution/Results: Tokenizer choice critically affects predictive performance; certain intrinsic metrics (e.g., opcode coverage, mnemonic preservation) correlate strongly with downstream accuracy. We propose a lightweight, binary-code-oriented tokenization optimization pathway that enhances semantic capture without compromising computational efficiency. This work establishes a reproducible evaluation paradigm and practical tokenizer design guidelines for adapting large language models to low-level code.
This study addresses the tokenization challenge for morphologically rich, low-resource languages (e.g., Turkish). We propose the first multidimensional evaluation framework tailored to such languages, built upon the newly constructed TR-MMLU dataset. We systematically benchmark mainstream tokenizers across four dimensions: vocabulary size, token count, processing efficiency, and preservation of linguistic structure—introducing two novel metrics: language-specific tokenization ratio (%TR) and token purity (%Pure). Results show that %TR strongly correlates with downstream task performance, whereas %Pure exhibits weak correlation; scaling model parameters alone does not improve linguistic understanding. Crucially, this work provides the first empirical validation that %TR is a more effective indicator of tokenization quality than conventional metrics. It underscores the necessity of linguistically grounded, language-specific tokenization strategies for morphologically complex languages and establishes a reproducible evaluation paradigm and methodological foundation for low-resource NLP.
Existing stemmer evaluation methods fail to quantify the semantic degradation caused by over-stemming in downstream tasks. This paper proposes the first task-oriented text normalization evaluation framework, overcoming the limitations of traditional lemmatization-based metrics by jointly measuring three dimensions: Stemming Effectiveness Score (SES), Model Performance Delta (MPD), and Average Normalized Levenshtein Distance (ANLD). The framework enables integrated analysis of both efficiency and semantic safety—the first of its kind. Empirical evaluation across multiple languages reveals a critical insight: high stemming recall does not necessarily improve downstream performance. For instance, the Bangla stemmer suffers performance degradation due to aggressive over-stemming, whereas the English Snowball stemmer achieves a superior trade-off between effectiveness and semantic fidelity. This work establishes a principled, task-aware methodology for evaluating and comparing stemmers beyond surface-form matching.
This work identifies and systematically characterizes the phenomenon of “intermediate merge residues” in Byte Pair Encoding (BPE) tokenizers—low-frequency subword tokens generated during training that are rarely used during inference. These redundant tokens unnecessarily consume vocabulary capacity and degrade model robustness to input noise and spelling errors. To address this issue, the authors propose LiteToken, a lightweight post-training optimization method that directly prunes existing BPE vocabularies without requiring retraining. By analyzing token formation mechanisms and empirical usage frequencies, LiteToken effectively removes superfluous tokens, substantially reducing vocabulary fragmentation and model parameter count while preserving original task performance and enhancing robustness to anomalous inputs.