Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review

📅 2025-10-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Pre-trained language models (PLMs) often underperform in domain-specific text classification due to domain-specific terminology, syntactic idiosyncrasies, and class imbalance. Method: We conduct a systematic literature review (2018–early 2024) of 41 studies, adhering to PRISMA guidelines and augmented by AI-assisted tools for rigorous screening; we propose the first taxonomy of PLM adaptation techniques for domain text classification and establish a cross-domain, multi-dimensional performance evaluation framework. Contribution/Results: Empirical analysis of Transformer-based models—including BERT, SciBERT, and BioBERT—across biomedical and other domains identifies domain adaptation and data bias as critical bottlenecks. We validate the efficacy of fine-tuning strategies, domain-aware self-supervised pre-training, and balanced sampling techniques. This work provides both theoretical foundations and practical guidelines for designing domain-adaptive PLMs, advancing reproducible and robust domain-specific NLP.

Technology Category

Machine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: Text Classification & Sentiment AnalysisApplication Domains: Humanities & Computational Social Science

Application Category

Search and Retrieval-Augmented AI: Vertical and domain-specific searchWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural language processing (NLP) plays a crucial role in addressing this challenge, particularly in text classification tasks. While large language models (LLMs) have achieved remarkable success in NLP, their accuracy can suffer in domain-specific contexts due to specialized vocabulary, unique grammatical structures, and imbalanced data distributions. In this systematic literature review (SLR), we investigate the utilization of pre-trained language models (PLMs) for domain-specific text classification. We systematically review 41 articles published between 2018 and January 2024, adhering to the PRISMA statement (preferred reporting items for systematic reviews and meta-analyses). This review methodology involved rigorous inclusion criteria and a multi-step selection process employing AI-powered tools. We delve into the evolution of text classification techniques and differentiate between traditional and modern approaches. We emphasize transformer-based models and explore the challenges and considerations associated with using LLMs for domain-specific text classification. Furthermore, we categorize existing research based on various PLMs and propose a taxonomy of techniques used in the field. To validate our findings, we conducted a comparative experiment involving BERT, SciBERT, and BioBERT in biomedical sentence classification. Finally, we present a comparative study on the performance of LLMs in text classification tasks across different domains. In addition, we examine recent advancements in PLMs for domain-specific text classification and offer insights into future directions and limitations in this rapidly evolving domain.
Problem

Research questions and friction points this paper is trying to address.

Addressing LLM accuracy issues in domain-specific text classification tasks
Investigating pre-trained language models for specialized vocabulary and structures
Systematically reviewing PLM applications across domains with performance comparisons
Innovation

Methods, ideas, or system contributions that make the work stand out.

Systematic review of pre-trained language models
Comparative experiment with BERT variants
Taxonomy of techniques for domain adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhyar Rzgar K. Rostam
Doctoral School of Applied Informatics and Applied Mathematics, Obuda University, Budapest, Hungary
G
Gábor Kertész
John von Neumann Faculty of Informatics, Obuda University, Budapest, Hungary and Laboratory of Parallel and Distributed Systems, Institute for Computer Science and Control (SZTAKI), Hungarian Research Network (HUN-REN), Budapest, Hungary