Leveraging Textual-Cues for Enhancing Multimodal Sentiment Analysis by Object Recognition

📅 2026-01-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal sentiment analysis faces significant challenges due to the substantial disparity between visual and textual modalities, ambiguous emotional expressions, and complex contextual semantics. To address these issues, this work proposes a Text-Enhanced Multimodal Sentiment Analysis (TEMSA) framework that, for the first time, explicitly incorporates the names of all detected objects in an image as supplementary textual cues. These object-derived terms are fused with the original input text to construct an enriched multimodal representation, effectively bridging the semantic gap between modalities. Extensive experiments on two standard benchmark datasets demonstrate that the proposed approach substantially outperforms both unimodal baselines and existing multimodal fusion strategies, confirming that comprehensive integration of object-level semantic information significantly enhances sentiment classification accuracy.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Multimodal sentiment analysis, which includes both image and text data, presents several challenges due to the dissimilarities in the modalities of text and image, the ambiguity of sentiment, and the complexities of contextual meaning. In this work, we experiment with finding the sentiments of image and text data, individually and in combination, on two datasets. Part of the approach introduces the novel `Textual-Cues for Enhancing Multimodal Sentiment Analysis'(TEMSA) based on object recognition methods to address the difficulties in multimodal sentiment analysis. Specifically, we extract the names of all objects detected in an image and combine them with associated text; we call this combination of text and image data TEMS. Our results demonstrate that only TEMS improves the results when considering all the object names for the overall sentiment of multimodal data compared to individual analysis. This research contributes to advancing multimodal sentiment analysis and offers insights into the efficacy of TEMSA in combining image and text data for multimodal sentiment analysis.
Problem

Research questions and friction points this paper is trying to address.

multimodal sentiment analysis
text-image fusion
sentiment ambiguity
contextual meaning
modality dissimilarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Sentiment Analysis
Object Recognition
Textual Cues
TEMSA
Image-Text Fusion
🔎 Similar Papers
No similar papers found.
Sumana Biswas
Sumana Biswas
Research Associate
Path planningMission planningAutonomous systemArtificial neural networkProduct management
K
Karen Young
School of Computer Science, University of Galway, Ireland
J
Josephine Griffith
School of Computer Science, University of Galway, Ireland