Multi-Level Attention and Contrastive Learning for Enhanced Text Classification with an Optimized Transformer

📅 2025-01-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient deep semantic modeling, low computational efficiency, and poor generalization under class imbalance or cross-domain settings in text classification, this paper proposes a lightweight Transformer architecture integrating multi-level attention and contrastive learning. Methodologically, it introduces a global-local collaborative multi-level attention mechanism to enhance fine-grained semantic capture; incorporates contrastive learning throughout both pretraining and fine-tuning stages to strengthen class discriminability; and employs a low-overhead feature projection module to significantly reduce computational redundancy. Experimental results on multiple benchmark datasets demonstrate that the model consistently outperforms BiLSTM, CNN, standard Transformer, and BERT in accuracy, F1-score, and recall. Moreover, it achieves substantial improvements in semantic representation quality and cross-domain generalization capability.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Transfer, Domain Adaptation, Multi-Task LearningComputer Vision: Multi-modal Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
This paper studies a text classification algorithm based on an improved Transformer to improve the performance and efficiency of the model in text classification tasks. Aiming at the shortcomings of the traditional Transformer model in capturing deep semantic relationships and optimizing computational complexity, this paper introduces a multi-level attention mechanism and a contrastive learning strategy. The multi-level attention mechanism effectively models the global semantics and local features in the text by combining global attention with local attention; the contrastive learning strategy enhances the model's ability to distinguish between different categories by constructing positive and negative sample pairs while improving the classification effect. In addition, in order to improve the training and inference efficiency of the model on large-scale text data, this paper designs a lightweight module to optimize the feature transformation process and reduce the computational cost. Experimental results on the dataset show that the improved Transformer model outperforms the comparative models such as BiLSTM, CNN, standard Transformer, and BERT in terms of classification accuracy, F1 score, and recall rate, showing stronger semantic representation ability and generalization performance. The method proposed in this paper provides a new idea for algorithm optimization in the field of text classification and has good application potential and practical value. Future work will focus on studying the performance of this model in multi-category imbalanced datasets and cross-domain tasks and explore the integration wi
Problem

Research questions and friction points this paper is trying to address.

Text Classification
Deep Semantic Understanding
Imbalanced Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Enhanced Transformer
Multi-layer Attention
Contrastive Learning
🔎 Similar Papers
No similar papers found.
Jia Gao
Jia Gao
Stevens Institute of Technology, Hoboken, USA
G
Guiran Liu
San Francisco State University, San Francisco, USA
B
Binrong Zhu
San Francisco State University, San Francisco, USA
Shicheng Zhou
Shicheng Zhou
Unknown affiliation
H
Hongye Zheng
The Chinese University of Hong Kong, Hong Kong, China
X
Xiaoxuan Liao
New York University, New York, USA