Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models

๐Ÿ“… 2025-02-18
๐Ÿ›๏ธ Journal of Advances in Information Technology
๐Ÿ“ˆ Citations: 30
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This paper addresses the poor generalization of sentiment analysis models to unseen emojis. Methodologically, it systematically quantifies the generalization gap of Universal Sentence Encoder (USE) and Sentence-BERT (SBERT) under in-distribution text versus out-of-distribution emoji shifts; integrates sentence embeddings with LSTM/fine-tuning heads; employs PyTorch Distributed Data Parallel (DDP) for scalable training; and incorporates SHAP for feature-level interpretability. Key contributions include: (1) the first quantitative characterization of performance degradation induced by emoji distribution shift; (2) a novel dual-path framework unifying distributed training and post-hoc explanation; (3) achieving 70% cross-emoji generalization accuracyโ€”matching 98% of baseline in-distribution accuracy; (4) 15% training speedup without accuracy loss; and (5) SHAP-based identification of emoji usage bias and salient sentiment-driving features.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
๐Ÿ“ Abstract
Emojis are being frequently used in todays digital world to express from simple to complex thoughts more than ever before. Hence, they are also being used in sentiment analysis and targeted marketing campaigns. In this work, we performed sentiment analysis of Tweets as well as on emoji dataset from the Kaggle. Since tweets are sentences we have used Universal Sentence Encoder (USE) and Sentence Bidirectional Encoder Representations from Transformers (SBERT) end-to-end sentence embedding models to generate the embeddings which are used to train the Standard fully connected Neural Networks (NN), and LSTM NN models. We observe the text classification accuracy was almost the same for both the models around 98 percent. On the contrary, when the validation set was built using emojis that were not present in the training set then the accuracy of both the models reduced drastically to 70 percent. In addition, the models were also trained using the distributed training approach instead of a traditional singlethreaded model for better scalability. Using the distributed training approach, we were able to reduce the run-time by roughly 15% without compromising on accuracy. Finally, as part of explainable AI the Shap algorithm was used to explain the model behaviour and check for model biases for the given feature set.
Problem

Research questions and friction points this paper is trying to address.

Evaluate sentiment analysis on text and emoji data
Compare end-to-end, transfer learning, and distributed AI models
Investigate explainable AI to understand model behavior and biases
Innovation

Methods, ideas, or system contributions that make the work stand out.

Used USE and SBERT for embeddings
Employed distributed training approach
Applied Shap for explainable AI
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
CR Rao AIMSCS | Independent Researcher