🤖 AI Summary
This paper addresses the technological evolution of social media sentiment analysis systems in the SemEval competition. Method: It systematically reviews the winning approaches from 2013 to 2021, comprehensively analyzing methodological shifts across the full pipeline—data acquisition, preprocessing, and classification—and comparatively evaluating lexicon-based methods, traditional machine learning, word embeddings (Word2Vec/GloVe), LSTM/CNN architectures, and Transformer-based models (e.g., BERT). The analysis synthesizes technical strategies from 658 participating teams. Contribution/Results: The study identifies a fundamental paradigm shift—from feature-engineering-driven pipelines to end-to-end fine-tuning of pretrained Transformers. Empirical evidence demonstrates that neural architectures, particularly pretrained Transformers, substantially improve accuracy and robustness. These findings provide empirical support and methodological guidance for rapid prototyping and next-generation competition system design.
📝 Abstract
<div class="page" title="Page 1"><div class="layoutArea"><div class="column"><div class="page" title="Page 1"><div class="layoutArea"><div class="column"><p>ocial media platforms are becoming the foundations of social interactions including messaging and opinion expression. In this regard, sentiment analysis techniques focus on providing solutions to ensure the retrieval and analysis of generated data including sentiments, emotions, and discussed topics. International competitions such as the International Workshop on Semantic Evaluation (SemEval) have attracted many researchers and practitioners with a special research interest in building sentiment analysis systems. In our work, we study top-ranking systems for each SemEval edition during the 2013-2021 period, a total of 658 teams participated in these editions with increasing interest over years. We analyze the proposed systems marking the evolution of research trends with a focus on the main components of sentiment analysis systems including data acquisition, preprocessing, and classification. Our study shows an active use of preprocessing techniques, an evolution of features engineering and word representation from lexicon-based approaches to word embeddings, and the dominance of neural networks and transformers over the classification phasefostering the use of ready-to-use models. Moreover, we provide researchers with insights based on experimented systems which will allow rapid prototyping of new systems and help practitioners build for future SemEval editions.</p></div></div></div></div></div></div>