A Retail-Corpus for Aspect-Based Sentiment Analysis with Large Language Models

📅 2025-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses multilingual aspect-based sentiment analysis (ABSA) in实体 retail settings. We introduce and publicly release the first large-scale, manually annotated multilingual customer review dataset for retail—comprising 10,814 samples across eight aspect categories with corresponding sentiment polarities—filling a critical gap in multilingual ABSA benchmark resources for the retail domain. Leveraging this dataset, we systematically evaluate GPT-4 and LLaMA-3 on the joint task of aspect identification and sentiment classification, employing domain-specific prompt engineering and rigorous experimental validation using fine-grained retail corpora. Results show both models achieve accuracy above 85%, with GPT-4 significantly outperforming LLaMA-3 across all key metrics. Our contribution includes: (1) a high-quality, multilingual, retail-specific ABSA benchmark dataset; and (2) an empirically grounded performance baseline for large language models in vertical-domain multilingual ABSA tasks.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Machine Translation, Multilinguality, Cross-Lingual NLPData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
Aspect-based sentiment analysis enhances sentiment detection by associating it with specific aspects, offering deeper insights than traditional sentiment analysis. This study introduces a manually annotated dataset of 10,814 multilingual customer reviews covering brick-and-mortar retail stores, labeled with eight aspect categories and their sentiment. Using this dataset, the performance of GPT-4 and LLaMA-3 in aspect based sentiment analysis is evaluated to establish a baseline for the newly introduced data. The results show both models achieving over 85% accuracy, while GPT-4 outperforms LLaMA-3 overall with regard to all relevant metrics.
Problem

Research questions and friction points this paper is trying to address.

Introduces annotated multilingual retail review dataset for aspect sentiment analysis
Evaluates GPT-4 and LLaMA-3 performance on aspect-based sentiment tasks
Establishes baseline accuracy exceeding 85% for retail sentiment analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Manually annotated multilingual retail dataset creation
Evaluated GPT-4 and LLaMA-3 performance comparison
Established over 85% accuracy baseline benchmarks
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Oleg Silcenco
University of Twente
M
Marcos R. Machado
University of Twente
Wallace C. Ugulino
Wallace C. Ugulino
University of Twente
D
Daniel Braun
Marburg University