RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

📅 2026-04-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of cross-lingual and cross-domain generalization in sentiment analysis for Romanian and Italian, where labeled data is scarce. To this end, the authors introduce the first large-scale multilingual, multi-domain sentiment analysis benchmark dataset, comprising 36,000 annotated reviews and over 200,000 unlabeled samples. They propose a multi-objective adversarial training framework that integrates meta-learning to dynamically balance the loss weights between sentiment discrimination and language- or domain-invariant representation learning, further enhanced by few-shot prompting. Experimental results demonstrate that a fine-tuned XLM-R model achieves an F1 score of 66.23%, outperforming baseline methods by 4.64%. Meanwhile, Llama-3.1-8B under a few-shot setting attains an F1 of 58.43%, effectively illustrating the trade-off between fine-tuning and prompt-based strategies.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Machine Translation, Multilinguality, Cross-Lingual NLPSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141 unlabeled samples. To address cross-lingual and cross-domain challenges, we propose a multi-target adversarial training framework that employs loss reversal with meta-learned coefficients to dynamically balance sentiment discrimination with domain and language invariance. XLM-R achieves an F1-score of 66.23% with our approach, outperforming the baseline by 4.64%. Few-shot evaluation shows that Llama-3.1-8B achieves 58.43% F1-score, revealing a meaningful trade-off between the efficiency of prompting-based approaches and the higher performance of task-specific fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

multilingual sentiment analysis
cross-lingual
cross-domain
Romanian
Italian
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial training
cross-lingual sentiment analysis
multi-domain adaptation
meta-learned coefficients
loss reversal
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Andrei-Marius Avram
Andrei-Marius Avram
Adobe
machine learningnatural language processingspeech recognition
A
Aureliu Valentin Antonie
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
C
Cosmin-Mircea Croitoru
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
V
Vlad Andrei Muntean
National University of Science and Technology POLITEHNICA Bucharest, Bucharest, Romania
Dumitru-Clementin Cercel
Dumitru-Clementin Cercel
Teaching Assistant of Computer Science, University Politehnica of Bucharest
Social Network AnalysisNatural Language ProcessingInformation RetrievalMachine Learning