Ensembling Multilingual Transformers for Robust Sentiment Analysis of Tweets

📅 2025-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the performance degradation of multilingual tweet sentiment analysis in low-resource languages due to scarce annotated data, this paper proposes an ensemble framework integrating multilingual BERT and XLM-R, augmented by large language models to enhance cross-lingual transfer capability. The method operates in a zero-shot setting—requiring no labeled data in target languages—and leverages model-level ensembling, joint multilingual training, and semantic alignment optimization to improve robustness in sentiment polarity classification. Evaluated on a multilingual tweet benchmark, the approach achieves 86.2% accuracy, outperforming individual base models by 3.7–5.1 percentage points, with particularly strong generalization to resource-scarce languages. This work delivers a scalable, low-dependency solution for unsupervised cross-lingual sentiment analysis.

Technology Category

Natural Language Processing: Machine Translation, Multilinguality, Cross-Lingual NLPMachine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Knowledge Representation Languages

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Sentiment analysis is a very important natural language processing activity in which one identifies the polarity of a text, whether it conveys positive, negative, or neutral sentiment. Along with the growth of social media and the Internet, the significance of sentiment analysis has grown across numerous industries such as marketing, politics, and customer service. Sentiment analysis is flawed, however, when applied to foreign languages, particularly when there is no labelled data to train models upon. In this study, we present a transformer ensemble model and a large language model (LLM) that employs sentiment analysis of other languages. We used multi languages dataset. Sentiment was then assessed for sentences using an ensemble of pre-trained sentiment analysis models: bert-base-multilingual-uncased-sentiment, and XLM-R. Our experimental results indicated that sentiment analysis performance was more than 86% using the proposed method.
Problem

Research questions and friction points this paper is trying to address.

Ensembling transformers for multilingual tweet sentiment analysis
Addressing sentiment analysis flaws in foreign languages
Improving performance without labeled training data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ensembling multilingual transformer models for sentiment analysis
Using pre-trained BERT and XLM-R models without labeled data
Achieving over 86% accuracy across multiple languages
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Meysam Shirdel Bilehsavar
Department of Computer Science, University of South Carolina, USA; Artificial Intelligence Institute, University of South Carolina, USA
Negin Mahmoudi
Negin Mahmoudi
Stevens Institute of Technology
Machine Learning
M
Mohammad Jalili Torkamani
School of Computing, University of Nebraska–Lincoln, Lincoln, Nebraska, USA
Kiana Kiashemshaki
Kiana Kiashemshaki
Bowling Green State University
Computer Science