Score
Designs and implements models, datasets, and processing pipelines that detect, classify, and quantify subjective opinion or emotional polarity in text. This includes document‑, sentence‑, and aspect‑level polarity or intensity labeling, choice of features or embeddings, supervised/unsupervised classifiers or regressors, and appropriate evaluation metrics and error analysis for sentiment tasks.
This paper addresses the technological evolution of social media sentiment analysis systems in the SemEval competition. Method: It systematically reviews the winning approaches from 2013 to 2021, comprehensively analyzing methodological shifts across the full pipeline—data acquisition, preprocessing, and classification—and comparatively evaluating lexicon-based methods, traditional machine learning, word embeddings (Word2Vec/GloVe), LSTM/CNN architectures, and Transformer-based models (e.g., BERT). The analysis synthesizes technical strategies from 658 participating teams. Contribution/Results: The study identifies a fundamental paradigm shift—from feature-engineering-driven pipelines to end-to-end fine-tuning of pretrained Transformers. Empirical evidence demonstrates that neural architectures, particularly pretrained Transformers, substantially improve accuracy and robustness. These findings provide empirical support and methodological guidance for rapid prototyping and next-generation competition system design.
Existing sentiment analysis of mobile app reviews predominantly focuses on coarse-grained polarity (e.g., positive/negative), lacking systematic modeling of fine-grained emotions (e.g., joy, anger, fear). Method: This work adapts Plutchik’s emotion wheel to the app review domain, establishing a structured annotation schema and releasing the first high-quality, human-annotated dataset for fine-grained emotion classification. We propose a hybrid annotation framework integrating iterative human labeling, large language model (LLM)-based automated labeling, and consistency evaluation. Contribution/Results: Experiments show strong agreement between LLM and human annotations (Cohen’s κ > 0.7), substantially reducing annotation cost. However, we also identify inherent limitations of fully automated labeling in complex linguistic contexts. To our knowledge, this is the first study addressing fine-grained emotion recognition in mobile app reviews, bridging a critical gap in user affect understanding within app ecosystems and establishing a new methodological paradigm.
Subjective language understanding—encompassing sentiment analysis, emotion recognition, sarcasm detection, humor comprehension, stance detection, metaphor interpretation, intent detection, and aesthetic assessment—faces inherent challenges including ambiguity, strong contextual dependency, and severe annotation scarcity. Method: We propose the first taxonomy for subjective language understanding tailored to large language models (LLMs), grounded in cognitive linguistics; develop a unified modeling framework integrating architectural analysis, cross-task comparative evaluation, and transfer learning experiments. Contribution/Results: Our taxonomy reveals shared cognitive mechanisms and modeling commonalities across tasks. Empirical validation confirms the framework’s effectiveness in enhancing generalization and interpretability. We systematically survey benchmark datasets and state-of-the-art methods, identify critical open issues—including model biases, data skewness, and ethical risks—and provide theoretical foundations and practical guidelines for developing explainable, robust, and human-centered subjective language processing systems.
To address the scarcity of cross-domain annotations, poor generalizability, and low reproducibility of supervised methods in aspect-category sentiment analysis (ACSA), this paper proposes a zero-shot large language model (LLM) framework. Methodologically, it integrates multiple chain-of-thought agents and introduces, for the first time, a token-level uncertainty quantification mechanism—dynamically weighting agent outputs via uncertainty scores. Combined with chain-of-thought prompting and Llama/Qwen models (3B–70B+), it achieves fine-grained sentiment classification without labeled data. Contributions include: (1) the first application of token-level uncertainty to assess decision reliability in zero-shot sentiment classification, significantly mitigating performance degradation under domain shift; and (2) empirical validation across model scales demonstrating concurrent improvements in prediction stability and accuracy—establishing a robust, reproducible solution for low-resource ACSA.
This study investigates language-specific polarity biases in multilingual sentiment analysis, revealing significant and directionally divergent disparities in how AI models classify positive and negative reviews across languages. Through a systematic comparison of encoder-based models and large language models on multilingual product review datasets, the work uncovers a pronounced negative bias in French large language models and a positive bias in Japanese encoder models—attributed to culturally embedded indirect criticism in Japanese discourse. These findings demonstrate that linguistic structures and cultural context profoundly shape model behavior, offering critical empirical evidence and cautionary insights for deploying sentiment analysis systems in multilingual commercial and societal applications.
Aspect-Sentiment-Opinion Triplet Extraction (ASTE) lacks annotated resources for Slavic languages, particularly Polish, which has no publicly available dataset. Method: We introduce the first Polish ASTE dataset, covering two domains—hotels and e-commerce—and strictly adhering to the standard English ASTE format to ensure cross-lingual comparability. The dataset is manually annotated with fine-grained sentiment structures and released under a CC-BY-NC license. Contribution/Results: Using this resource, we conduct the first systematic evaluation of two mainstream ASTE paradigms and two Polish large language models, revealing critical performance bottlenecks of existing methods on Slavic languages. This work fills a key gap in low-resource, fine-grained sentiment analysis for Slavic languages and establishes a benchmark dataset and empirical foundation for future multilingual ASTE research and model development.
Traditional sentiment analysis relies on discrete classification, which struggles to capture the nuanced gradations of emotional intensity required in domains such as finance. This work proposes a novel paradigm that reframes sentiment analysis as a continuous regression task by constructing a dataset annotated with fine-grained emotion intensity scores and fine-tuning open-source generative language models to predict values on a 0–100 scale. The proposed approach significantly outperforms conventional classification baselines and demonstrates strong cross-construct transferability on related affective tasks, including sentiment polarity and arousal. By enabling more precise and expressive modeling of emotional intensity, this method offers enhanced practical utility for real-world applications demanding granular affective understanding.
This study addresses the challenges of contextual understanding and implicit sentiment recognition in movie review sentiment classification by systematically evaluating a range of approaches, from traditional statistical models—such as TF-IDF combined with Naive Bayes, Logistic Regression, SVM, and LightGBM—to modern deep learning architectures, including LSTM, DistilBERT, and RoBERTa. Building upon this comparative analysis, the authors propose a soft voting–based ensemble model to leverage the complementary strengths of individual classifiers. Experimental results on the IMDb dataset demonstrate that the RoBERTa single model achieves an accuracy of 93.02%, while the ensemble model further enhances overall performance, yielding superior results in terms of both accuracy and F1 score. These findings confirm the effectiveness and robustness of ensemble strategies for sentiment analysis tasks.
This study addresses the challenge of scarce labeled data in sentiment analysis for software engineering, where off-the-shelf sentiment analysis tools often underperform. It presents the first systematic evaluation of various zero-shot learning (ZSL) approaches—including embedding-based methods, natural language inference, TARS, and generative models—on this task, leveraging an expert-defined sentiment label taxonomy. The authors compare these ZSL methods against fine-tuned Transformer models under diverse label settings. Experimental results demonstrate that certain ZSL approaches achieve macro F1 scores comparable to those of fine-tuned models, substantially reducing reliance on annotated data. Error analysis further reveals that subjective labeling practices and confusion between polarity and factual content are primary sources of misclassification.
Existing research on affective polarization is constrained by the scarcity of real-world data, high subjectivity in annotation, and the absence of a unified computational framework. This work proposes the first large language model (LLM)-driven multi-agent simulation platform that leverages natural language to generate context-aware virtual users, enabling the modeling of complex social interactions and emotional dynamics within synthetic social environments. The framework supports configurable, multi-level polarization scenarios, substantially enhancing reproducibility and cross-study comparability. Through experiments, the platform successfully replicates and extends key polarization phenomena documented in social psychology, demonstrating its effectiveness and potential for efficiently and flexibly investigating the dynamic mechanisms underlying affective polarization.
This study addresses polarization in multilingual texts by proposing a prompt engineering–based approach for fine-grained detection, encompassing three subtasks: binary polarization classification, polarization type identification, and manifestation categorization. The authors systematically design twelve prompt templates, varying along dimensions such as term clarity, definition specificity, reasoning guidance, and contextual exemplars, thereby offering the first comprehensive evaluation of how prompt design influences multilingual polarization detection. Experiments are conducted on the Aya-101 and Gemma3-27B models, with the latter achieving average macro F1 scores of 0.762, 0.587, and 0.444 across 22 languages. These results delineate both the capabilities and limitations of prompt engineering in coarse- and fine-grained sociolinguistic classification tasks.