stopword removal

Designs, implements, and evaluates text preprocessing components that identify and remove or otherwise handle stopwords from token streams, including curated and data-driven stopword lists, context-sensitive filtering, and normalization-aware token decisions. Measures and analyzes the effect of stopword choices on downstream NLP tasks and pipelines, and builds tools to configure, optimize, and audit stopword handling strategies.

stopwordremoval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic investigation into the ordering of preprocessing steps in sentiment analysis, which has constrained model performance and efficiency. Focusing on Twitter data, it presents the first quantitative evaluation of the relative impact and optimal sequencing of key preprocessing techniques—including tokenization, text cleaning, stemming, stopword removal (with negations preserved), and spelling correction. Through comprehensive combinatorial experiments, the work demonstrates that tokenization contributes most significantly to model performance, while spelling correction has the least effect. The identified optimal pipeline—tokenization followed by text cleaning, stemming, and stopword removal—substantially enhances model effectiveness and reduces trial-and-error costs, establishing a reproducible and efficient preprocessing paradigm for sentiment analysis.

machine learningpreprocessingsentiment analysis

Investigating Large Language Models' Linguistic Abilities for Text Preprocessing

Oct 13, 2025
MB
Marco Braga
🏛️ University of Milano-Bicocca | Politecnico di Torino

Conventional text preprocessing techniques—such as stopword removal, lemmatization, and stemming—rely heavily on language-specific linguistic rules and ignore contextual information, limiting their generalizability across multilingual settings. Method: This paper pioneers a systematic investigation of large language models (LLMs) as context-aware, universal preprocessors. Leveraging prompt engineering, we uniformly perform the three preprocessing tasks across six European languages without language-specific annotations or handcrafted rules. Contribution/Results: Experiments show LLMs achieve 97%, 82%, and 74% accuracy on stopword removal, lemmatization, and stemming, respectively. Downstream text classification models fed with LLM-preprocessed inputs attain up to a 6-percentage-point improvement in F1 score. This work demonstrates the feasibility and effectiveness of LLM-driven, end-to-end, context-sensitive, and multilingual-compatible text preprocessing—establishing a novel paradigm that reduces reliance on manual linguistic rules and enhances preprocessing robustness.

LLMs address context-dependent text preprocessing limitationsLLMs improve text classification accuracy over traditional techniquesTraditional methods ignore contextual information in preprocessing

How Does A Text Preprocessing Pipeline Affect Ontology Syntactic Matching?

Nov 06, 2024
ZQ
Zhangcheng Qiang
🏛️ Australian National University | Monash University

Non-standardized text preprocessing leads to unstable ontology matching (OM) results. Method: We systematically evaluate the impact of a four-stage preprocessing pipeline—tokenization, normalization, stopword removal, and stemming/lemmatization—across 49 OAEI ontology alignment tasks. Our analysis reveals that the first stage (tokenization and normalization) contributes significantly more to matching performance than subsequent stages. Based on this finding, we propose a context-driven preprocessing repair mechanism: (i) dynamically constructing a preservation lexicon to suppress spurious alignments; and (ii) integrating large language models (LLMs) into the pipeline via function calling—bypassing prompt engineering to prevent ground-truth drift. Contribution/Results: Evaluated across eight OAEI tracks, our approach substantially improves matching accuracy and demonstrates robustness. It establishes a reproducible, structured paradigm for synergistic integration of LLMs with classical NLP preprocessing in ontology matching.

Investigates text preprocessing impact on ontology matching.Proposes context-based repair to improve mapping accuracy.Recommends integrating preprocessing with large language models.

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

What Are They Filtering Out? A Survey of Filtering Strategies for Harm Reduction in Pretraining Datasets

Feb 17, 2025
MS
M. Stranisci
🏛️ University of Turin | aequa-tech | IT University of Copenhagen

Pretraining data filtering strategies intended to reduce harmful content inadvertently exacerbate representational underrepresentation of marginalized groups, thereby amplifying demographic bias at the data level. Method: We systematically reviewed 55 English-language large language model technical reports to construct the first integrated data governance evaluation framework balancing safety and fairness. Through controlled experiments and quantitative bias analysis across mainstream filtering strategies, we measured their impact on group-level representation. Contribution/Results: Our analysis reveals that such strategies reduce text associated with disadvantaged groups by 12.7%–38.4% on average—significantly worsening representational disparity. This study provides the first empirical evidence refuting the “safety implies fairness” assumption in AI governance. We propose a co-optimization paradigm that jointly addresses content safety and equitable group representation, advocating for fairness-aware data curation in foundation model development.

Analyzing side effects of content removal on dataset diversityAssessing filtering impact on vulnerable groups' representationEvaluating data filtering strategies for harm reduction in LLMs

Latest Papers

What's happening recently
View more

This work identifies and systematically characterizes the phenomenon of “intermediate merge residues” in Byte Pair Encoding (BPE) tokenizers—low-frequency subword tokens generated during training that are rarely used during inference. These redundant tokens unnecessarily consume vocabulary capacity and degrade model robustness to input noise and spelling errors. To address this issue, the authors propose LiteToken, a lightweight post-training optimization method that directly prunes existing BPE vocabularies without requiring retraining. By analyzing token formation mechanisms and empirical usage frequencies, LiteToken effectively removes superfluous tokens, substantially reducing vocabulary fragmentation and model parameter count while preserving original task performance and enhancing robustness to anomalous inputs.

BPE tokenizationintermediate merge residuestoken fragmentation

This work addresses the limitations of conventional tokenizers—such as Byte Pair Encoding (BPE)—which often misalign with linguistic structure, exacerbate biases, and inefficiently consume model capacity in multilingual and multidomain settings, while lacking systematic design and evaluation protocols. Treating tokenization as a core modeling decision for large language models, this study proposes the first context-aware framework that co-designs tokenization with the model architecture, integrating linguistic knowledge, domain-specific characteristics, and deployment constraints. By jointly optimizing the tokenizer and the model, and establishing standardized evaluation benchmarks alongside transparent reporting practices, the research provides both theoretical foundations and practical pathways toward more equitable, efficient, and adaptable language technologies, significantly enhancing performance and robustness across diverse languages and domains.

bias amplificationlarge language modelslinguistic alignment

Current large language models still struggle with challenges in natural language–driven data preparation tasks, including ambiguous user intent, real-world dirty data, and the generation of interpretable workflows. To address this gap, this work proposes PrepBench, the first systematic benchmark specifically designed for data preparation, which evaluates three core capabilities: interactive disambiguation, data preparation code generation, and transformation of code into visual workflows. Built upon real-world, multi-domain datasets, PrepBench features complex operations spanning 3 to 18 steps and code tasks up to 300 lines long, thereby filling a critical void in existing code generation benchmarks that lack data preparation scenarios. Experimental results demonstrate that state-of-the-art models exhibit limited performance on these tasks, and PrepBench effectively uncovers key bottlenecks, offering a reliable evaluation standard for future research.

code generationdata preparationlarge language models

Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models

Dec 03, 2025
TP
Taido Purason
🏛️ University of Tartu | Technical University of Applied Sciences Würzburg-Schweinfurt

Pre-trained tokenizers suffer from inefficient vocabulary expansion and imprecise pruning of redundant tokens during cross-domain or cross-lingual transfer. Method: This paper proposes a dynamic vocabulary optimization framework based on continued Byte-Pair Encoding (BPE) training. It incrementally integrates new vocabulary by extending the original BPE merge process, thereby improving token utilization; additionally, it introduces, for the first time, a leaf-node pruning strategy grounded in the BPE tree structure, enabling controllable and interpretable vocabulary reduction without compromising model performance. Contribution/Results: Experiments across multilingual settings and model families (e.g., BERT, XLM-R) demonstrate an average 15% vocabulary compression, substantial improvements in tokenization efficiency, and a 32% increase in usage rate of newly added tokens—establishing a robust, efficient paradigm for tokenizer customization and adaptation.

Adapt tokenizers for new domains or languages efficientlyExtend vocabulary without creating unused tokensPrune redundant tokens while maintaining model quality

This work proposes a token-level filtering intervention to efficiently, robustly, and cost-effectively attenuate undesirable capabilities of language models in specific domains—such as healthcare—during pretraining. By integrating sparse autoencoders for fine-grained token annotation and distilling a lightweight classifier, the method precisely removes signals associated with the target domain at the token level rather than discarding entire documents. Experiments demonstrate that this approach achieves up to a 7,000-fold computational slowdown in the model’s target-domain capability on the largest evaluated model, substantially outperforming document-level filtering. Crucially, it preserves beneficial general capabilities, remains compatible with downstream alignment processes, and exhibits robustness under noisy labeling conditions.

capability shapinglanguage model safetypretraining data filtering

Hot Scholars

HW

Hassan Wasswa

University of New South Wales (UNSW)
Deep LearningInternet of ThingsCybersecurityComputer Vision
UT

Ugur Turhan

Assistant Professor (Senior Lecturer), Aviation, School of Science, UNSW Canberra
Air Traffic ManagementAirport OperationsHuman FactorsSafety Management in Aviation
AN

Aziida Nanyonga

University of New South Wales (UNSW), Australia
Artificial IntelligenceMachine learningDeep LearningNatural Language Processing
GW

Graham Wild

UNSW Canberra
acousticsaerospaceaviationsensors
OM

Oleksandra Molloy

Senior Lecturer in Aviation, University of New South Wales
Aviation safetyroad safetyhuman factorsdrones