🤖 AI Summary
This work addresses multilingual aspect-based sentiment analysis (ABSA) in实体 retail settings. We introduce and publicly release the first large-scale, manually annotated multilingual customer review dataset for retail—comprising 10,814 samples across eight aspect categories with corresponding sentiment polarities—filling a critical gap in multilingual ABSA benchmark resources for the retail domain. Leveraging this dataset, we systematically evaluate GPT-4 and LLaMA-3 on the joint task of aspect identification and sentiment classification, employing domain-specific prompt engineering and rigorous experimental validation using fine-grained retail corpora. Results show both models achieve accuracy above 85%, with GPT-4 significantly outperforming LLaMA-3 across all key metrics. Our contribution includes: (1) a high-quality, multilingual, retail-specific ABSA benchmark dataset; and (2) an empirically grounded performance baseline for large language models in vertical-domain multilingual ABSA tasks.
📝 Abstract
Aspect-based sentiment analysis enhances sentiment detection by associating it with specific aspects, offering deeper insights than traditional sentiment analysis. This study introduces a manually annotated dataset of 10,814 multilingual customer reviews covering brick-and-mortar retail stores, labeled with eight aspect categories and their sentiment. Using this dataset, the performance of GPT-4 and LLaMA-3 in aspect based sentiment analysis is evaluated to establish a baseline for the newly introduced data. The results show both models achieving over 85% accuracy, while GPT-4 outperforms LLaMA-3 overall with regard to all relevant metrics.