🤖 AI Summary
To address insufficient deep semantic modeling, low computational efficiency, and poor generalization under class imbalance or cross-domain settings in text classification, this paper proposes a lightweight Transformer architecture integrating multi-level attention and contrastive learning. Methodologically, it introduces a global-local collaborative multi-level attention mechanism to enhance fine-grained semantic capture; incorporates contrastive learning throughout both pretraining and fine-tuning stages to strengthen class discriminability; and employs a low-overhead feature projection module to significantly reduce computational redundancy. Experimental results on multiple benchmark datasets demonstrate that the model consistently outperforms BiLSTM, CNN, standard Transformer, and BERT in accuracy, F1-score, and recall. Moreover, it achieves substantial improvements in semantic representation quality and cross-domain generalization capability.
📝 Abstract
This paper studies a text classification algorithm based on an improved Transformer to improve the performance and efficiency of the model in text classification tasks. Aiming at the shortcomings of the traditional Transformer model in capturing deep semantic relationships and optimizing computational complexity, this paper introduces a multi-level attention mechanism and a contrastive learning strategy. The multi-level attention mechanism effectively models the global semantics and local features in the text by combining global attention with local attention; the contrastive learning strategy enhances the model's ability to distinguish between different categories by constructing positive and negative sample pairs while improving the classification effect. In addition, in order to improve the training and inference efficiency of the model on large-scale text data, this paper designs a lightweight module to optimize the feature transformation process and reduce the computational cost. Experimental results on the dataset show that the improved Transformer model outperforms the comparative models such as BiLSTM, CNN, standard Transformer, and BERT in terms of classification accuracy, F1 score, and recall rate, showing stronger semantic representation ability and generalization performance. The method proposed in this paper provides a new idea for algorithm optimization in the field of text classification and has good application potential and practical value. Future work will focus on studying the performance of this model in multi-category imbalanced datasets and cross-domain tasks and explore the integration wi