Score
Detecting and modeling expressed emotions or sentiment from modalities such as facial expressions or text—potentially as continuous signals—and relating those measurements to outcomes like engagement, coping styles, or anthropomorphization.
This paper addresses the core challenge of insufficient machine empathy in affective computing. We propose a unified framework integrating large language models (LLMs), multimodal learning (text, speech, and physiological signals), and personalized modeling. Through a systematic review, we analyze advances in emotion recognition, sentiment analysis, and personality modeling across four key application domains: AI chatbots, multimodal human–computer interaction, mental health interventions, and safety-critical systems—revealing empirical patterns linking data modality, scale, and diversity to model performance. We introduce, for the first time, a comprehensive research paradigm encompassing ethical assessment, annotated dataset analysis, and verifiability-oriented design, thereby clarifying technical trajectories and identifying critical research gaps. Finally, we formulate a tripartite design principle—“safety–empathy–utility”—for next-generation affective support systems, accompanied by an empirically grounded validation pathway.
This paper addresses the challenges of recognizing three complex psychological states—stress, depression, and engagement—and the lack of a unified modeling framework for their joint analysis. Methodologically, it conducts the first systematic review simultaneously covering all three states, integrating multimodal data (speech, text, physiological signals), feature- and decision-level fusion strategies, and both machine learning and deep learning models; it performs cross-benchmark comparative analysis of state-of-the-art performance on major datasets including DAIC-WOZ, AVEC, and RECOLA. Key contributions include: (1) proposing the first unified taxonomy and technology evolution timeline for these three states; (2) establishing a general computational analysis pipeline; and (3) identifying critical bottlenecks in model interpretability, cross-population generalizability, and privacy preservation. The work provides both theoretical foundations and practical guidelines for computational modeling of atypical psychological states.
Empathic response detection faces challenges including ill-defined task formulations, fragmented multimodal modeling, and the absence of systematic surveys. This paper conducts a cross-modal systematic review of 62 high-quality studies. Methodologically, it unifies network architecture design principles along modality dimensions for the first time, proposes a standardized three-tier interaction framework (individual, dyadic, and group) and a five-category task taxonomy—including local/global empathy recognition and emotional contagion detection. It constructs the first structured empathy detection knowledge graph, integrating four input modalities (text, audio, audiovisual, and physiological signals), twelve publicly available datasets, and seven reproducible codebases. By synergizing NLP, audiovisual modeling, speech emotion analysis, and time-frequency physiological signal processing—augmented with cross-modal contrastive learning and meta-analysis—the study identifies critical research gaps, establishing both theoretical foundations and practical guidelines for robust empathic computing.
Existing research often examines textual, visual, or audio modalities of short videos in isolation, failing to uncover how their interplay influences user engagement. This work proposes the first reproducible and interpretable multimodal analysis framework that integrates automated feature extraction with Shapley-value-based attribution to systematically investigate how multimodal interactions affect view counts in TikTok content related to social anxiety disorder. The study reveals that facial expressions are more predictive than textual sentiment, that informational content garners greater attention than emotional support, and that multimodal synergies exhibit strong threshold-dependent effects—thereby transcending the limitations of conventional unimodal analyses.
To address the limitations of conventional emotion recognition in human-computer interaction—namely, oversimplified modeling and insufficient coverage of discrete emotion categories—this paper proposes a multimodal continuous emotion modeling framework. It fuses facial expression, prosodic, and textual transcription features to construct a continuous emotion representation in the three-dimensional Valence-Arousal-Dominance (VAD) space. A novel contribution is the adaptive mapping of discrete emotion labels to the VAD space via K-means clustering, enabling open-vocabulary emotion generation. Additionally, a cross-modal feature alignment mechanism and a joint classifier are designed to enhance multimodal fusion. Evaluated on the MER2024 Chinese film-and-television dataset, the method achieves a 42% improvement in emotion vocabulary coverage and an accuracy of 91.3%, effectively balancing emotional diversity and discriminative precision.
This study addresses automated multimodal depression detection by proposing a cross-modal modeling framework integrating audio, visual, and textual features. Methodologically, it introduces a synergistic architecture combining XGBoost for shallow temporal statistical feature learning, Transformer networks for intra-modal long-range dependency modeling, and large language models (LLMs) for enhanced textual semantic and contextual understanding, augmented by modality alignment and weighted fusion strategies. Comprehensive evaluation on public multimodal depression datasets demonstrates that the proposed fusion model significantly outperforms unimodal baselines—achieving AUC improvements of 4.2–7.8 percentage points—thereby validating the complementary strengths of heterogeneous models and the efficacy of multimodal representation learning. The work establishes a novel, interpretable, and robust paradigm for mental health status assessment, offering a reproducible technical pathway grounded in principled multimodal integration.
This work investigates whether large language models (LLMs) can perform affective semantic understanding solely from valence-arousal (VA) numerical representations—without visual input. We evaluate zero-shot emotion classification and prompt-driven semantic description generation on the IIMI (basic emotions) and Emotic (complex emotions) datasets; VA features are extracted by FaceChannel, and models including GPT and Llama are assessed. Results show that LLMs exhibit limited performance on VA-based classification—particularly in distinguishing non-binary emotions—yet achieve strong performance in semantic description generation (BLEU-4 = 0.62; human-rated semantic similarity = 4.3/5), demonstrating robust mapping from abstract affective dimensions to natural language. To our knowledge, this is the first study to empirically validate that LLMs can directly interpret non-visual, structured affective representations. The findings establish a novel paradigm for lightweight, bias-mitigated affective computing grounded in interpretable, modality-agnostic emotional features.
This study addresses the substantial inter-annotator disagreement and low stability observed in sentiment annotation. Grounded in Appraisal Theory, it systematically investigates the textual features of sentiment expression and their computability. Using a manually annotated industrial corpus, we construct a fine-grained sentiment dataset and train language models to emulate human annotation behavior. Methodologically, we move beyond mainstream sentiment classification paradigms by incorporating appraisal-oriented semantic structures—namely, Attitude, Engagement, and Graduation—which remain underexplored in computational linguistics. This enables us to uncover stable statistical patterns and language-driven mechanisms underlying annotation discrepancies. Experiments demonstrate that our model effectively discriminates among distinct appraisal-based sentiment contexts, achieving significant improvements over baselines in cross-context generalization and fine-grained linguistic cue modeling. Our work establishes a theory-informed modeling framework for sentiment computation and provides an interpretable, appraisal-grounded evaluation benchmark.
This study addresses key challenges in multimodal emotion recognition—namely, the scarcity of affectively annotated data, semantic gaps across modalities, and opaque reasoning processes—by proposing a novel paradigm termed “MER-with-LLMs.” The work systematically constructs a large language model (LLM)-based framework that integrates three core components: affective data augmentation, cross-modal alignment, and interpretable reasoning. It establishes, for the first time, a comprehensive taxonomy and research roadmap for leveraging LLMs in general-purpose affective intelligence. Through a thorough survey of existing approaches, the paper delineates a clear academic landscape, elucidating the field’s developmental trajectory, critical bottlenecks, and emerging directions, thereby advancing multimodal emotion recognition toward greater structural coherence and interpretability.
This study investigates the alignment between large language models (LLMs) and human self-reported emotions in fine-grained sentiment recognition—moving beyond conventional coarse-grained classification. To this end, we introduce EXPRESS, the first benchmark dataset for fine-grained self-reported emotion analysis, comprising 251 authentic user-emotion labels. Grounded in classical emotion theory, we decompose model predictions into eight basic emotion categories, establishing an interpretable, fine-grained alignment evaluation framework. Through systematic experiments across diverse prompting strategies, we evaluate leading LLMs and find that while they generate theoretically plausible emotion terms, their capacity to capture context-dependent, nuanced affective states remains substantially inferior to human performance. Our core contributions are threefold: (1) the EXPRESS dataset; (2) a novel fine-grained emotion decomposition and alignment evaluation paradigm; and (3) empirical characterization of the fundamental limits of LLMs’ emotional alignment capability.
Existing methods struggle to jointly model the similarity and conflict between affective signals across modalities (e.g., image and text). To address this, we propose the first large language model (LLM)-based representation decomposition framework that explicitly disentangles cross-modal shared affective components from modality-specific cues. Our approach employs an attention-driven dynamic soft prompting mechanism for adaptive multimodal fusion and guides the fine-tuning of multimodal LLMs. The architecture integrates a pretrained multimodal encoder, a representation decomposition network, and a soft prompt generation module. Evaluated on three benchmark tasks—sentiment analysis, emotion recognition, and hate meme detection—our method consistently outperforms state-of-the-art approaches, achieving average accuracy gains of 2.3–4.1%. Notably, it is the first to jointly model both congruent and adversarial evidence in multimodal affective understanding.
Multimodal affective computing (MAC) suffers from unstable performance and insufficient understanding of how model architectures and data characteristics jointly influence performance. To address this, we propose a hybrid optimization framework integrating generative knowledge prompting, cross-modal alignment, and supervised fine-tuning, accompanied by a systematic benchmark for comprehensive evaluation of state-of-the-art open-source multimodal large language models (MLLMs) on audio-visual-text fusion-based emotion recognition. Extensive experiments across multiple standard benchmarks demonstrate significant improvements in end-to-end emotion analysis accuracy and robustness. This work provides the first empirical characterization of the synergistic interplay between architectural design choices and data properties in multimodal emotion understanding, establishing an interpretable and reproducible paradigm for MAC model development. The implementation is publicly available.
Existing sign language emotion recognition research is severely limited due to the dual functional role of facial expressions and manual signs—both encode grammatical structure and emotional content—leading to feature coupling and annotation ambiguity. Method: We introduce the first fine-grained annotated American Sign Language (ASL) emotion video dataset, comprising 200 multimodal samples, with collaborative labeling by three Deaf interpreters across six discrete emotions, sentiment polarity, and open-ended emotional cue descriptions. We systematically disentangle syntax- and emotion-shared facial and manual features, establishing the first ASL emotion recognition benchmark. Contribution/Results: We propose a ViT+LSTM+CLIP multimodal fusion baseline, achieving 58.3% accuracy on emotion classification—substantially outperforming random chance. The dataset is publicly released on Hugging Face, addressing a critical gap in sign language affective computing.