Natural Language Processing Models for Robust Document Categorization

πŸ“… 2026-02-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of balancing accuracy and computational efficiency in automated document classification under class-imbalanced conditions by systematically evaluating three representative models: Naive Bayes, bidirectional LSTM (BiLSTM), and fine-tuned BERT. Experimental results demonstrate that while BERT achieves over 99% accuracy, it incurs substantial computational overhead; Naive Bayes offers the fastest training speed but yields only approximately 94.5% accuracy; BiLSTM strikes the best trade-off between performance and efficiency, attaining 98.56% accuracy. Leveraging these insights, the authors implement a lightweight, robust demonstration system for automated technical request routing, confirming BiLSTM’s practical suitability for real-world deployment and offering an effective solution for text classification in resource-constrained environments.

Technology Category

Natural Language Processing: Text Classification & Sentiment AnalysisMachine Learning: Multi-class/Multi-label Learning & Extreme ClassificationData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
πŸ“ Abstract
This article presents an evaluation of several machine learning methods applied to automated text classification, alongside the design of a demonstrative system for unbalanced document categorization and distribution. The study focuses on balancing classification accuracy with computational efficiency, a key consideration when integrating AI into real world automation pipelines. Three models of varying complexity were examined: a Naive Bayes classifier, a bidirectional LSTM network, and a fine tuned transformer based BERT model. The experiments reveal substantial differences in performance. BERT achieved the highest accuracy, consistently exceeding 99\%, but required significantly longer training times and greater computational resources. The BiLSTM model provided a strong compromise, reaching approximately 98.56\% accuracy while maintaining moderate training costs and offering robust contextual understanding. Naive Bayes proved to be the fastest to train, on the order of milliseconds, yet delivered the lowest accuracy, averaging around 94.5\%. Class imbalance influenced all methods, particularly in the recognition of minority categories. A fully functional demonstrative system was implemented to validate practical applicability, enabling automated routing of technical requests with throughput unattainable through manual processing. The study concludes that BiLSTM offers the most balanced solution for the examined scenario, while also outlining opportunities for future improvements and further exploration of transformer architectures.
Problem

Research questions and friction points this paper is trying to address.

document categorization
class imbalance
computational efficiency
classification accuracy
natural language processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

document categorization
class imbalance
BiLSTM
BERT
computational efficiency
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
R
Radoslaw Roszczyk
Faculty of Electrical Engineering, Warsaw University of Technology
P
Pawel Tecza
Student of Warsaw University of Technology
M
Maciej Stodolski
Faculty of Electrical Engineering, Warsaw University of Technology
K
Krzysztof Siwek
Faculty of Electrical Engineering, Warsaw University of Technology