🤖 AI Summary
To address insufficient cross-modal interaction and limited fine-grained classification accuracy in multimodal sentiment analysis, this paper proposes a dynamic sentiment analysis model based on early feature fusion and a lightweight cross-modal attention mechanism. Methodologically, it adopts a Transformer architecture with a novel cross-modal interaction module tailored for multi-head attention, enabling joint encoding of textual, acoustic, and visual features at the input layer. It is the first work to systematically validate the significant superiority of early fusion on the CMU-MOSEI benchmark. Experiments show that the model achieves 72.39% accuracy—over 4 percentage points higher than representative late-fusion baselines—while reducing parameter count by 18%, thus balancing performance and efficiency. Key contributions are: (1) empirical confirmation of early fusion’s positive impact on multimodal sentiment representation; and (2) a low-overhead, highly compatible cross-modal attention strategy.
📝 Abstract
This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions between these modalities, thereby enabling more accurate and nuanced sentiment interpretation. The study evaluates three feature fusion strategies -- late stage fusion, early stage fusion, and multi-headed attention -- within a transformer-based architecture. Experiments were conducted using the CMU-MOSEI dataset, which includes synchronized text, audio, and visual inputs labeled with sentiment scores. Results show that early stage fusion significantly outperforms late stage fusion, achieving an accuracy of 71.87%, while the multi-headed attention approach offers marginal improvement, reaching 72.39%. The findings suggest that integrating modalities early in the process enhances sentiment classification, while attention mechanisms may have limited impact within the current framework. Future work will focus on refining feature fusion techniques, incorporating temporal data, and exploring dynamic feature weighting to further improve model performance.