Score
Designs, builds, and evaluates systems that detect, model, and generate human affective (emotional) states and responses; this work includes collecting and preprocessing multimodal signals (facial expression, speech, physiological measures, text), developing models for emotion recognition and response synthesis, and defining metrics and experiments to assess affective appropriateness, robustness, and safety.
This paper addresses the core challenge of insufficient machine empathy in affective computing. We propose a unified framework integrating large language models (LLMs), multimodal learning (text, speech, and physiological signals), and personalized modeling. Through a systematic review, we analyze advances in emotion recognition, sentiment analysis, and personality modeling across four key application domains: AI chatbots, multimodal human–computer interaction, mental health interventions, and safety-critical systems—revealing empirical patterns linking data modality, scale, and diversity to model performance. We introduce, for the first time, a comprehensive research paradigm encompassing ethical assessment, annotated dataset analysis, and verifiability-oriented design, thereby clarifying technical trajectories and identifying critical research gaps. Finally, we formulate a tripartite design principle—“safety–empathy–utility”—for next-generation affective support systems, accompanied by an empirically grounded validation pathway.
Physiological emotion data collection suffers from misalignment between subjective labels and objective physiological responses, primarily due to human participant dependency and associated cognitive biases. Method: We conducted a VR-based emotion elicitation study with 37 participants, complemented by semi-structured interviews, to investigate participant-centered factors affecting labeling fidelity. Contribution/Results: We identify three critical human-induced interference factors—perceptual bias, experimental design mismatch, and environmental mismatch—and provide the first systematic characterization of the decoupling mechanism between subjective cognition and physiological response. Based on these findings, we propose a participant-centered experimental design paradigm and a context-enhanced annotation framework, yielding seven actionable guidelines for physiological emotion data collection. This work establishes a human-factor foundation for reliable affective labeling and advances AI-driven affective modeling by bridging cognitive and physiological domains.
This paper addresses the challenges of recognizing three complex psychological states—stress, depression, and engagement—and the lack of a unified modeling framework for their joint analysis. Methodologically, it conducts the first systematic review simultaneously covering all three states, integrating multimodal data (speech, text, physiological signals), feature- and decision-level fusion strategies, and both machine learning and deep learning models; it performs cross-benchmark comparative analysis of state-of-the-art performance on major datasets including DAIC-WOZ, AVEC, and RECOLA. Key contributions include: (1) proposing the first unified taxonomy and technology evolution timeline for these three states; (2) establishing a general computational analysis pipeline; and (3) identifying critical bottlenecks in model interpretability, cross-population generalizability, and privacy preservation. The work provides both theoretical foundations and practical guidelines for computational modeling of atypical psychological states.
A lack of high-quality, multimodal benchmark datasets hinders progress in French affective computing. Method: This paper introduces FERG, the first French multimodal emotion dataset grounded in card-game interactions. It captures natural emotional expressions from 20 participants across 10 conversational gameplay sessions, synchronously recording facial video, speech, and hand-motion capture data. Emotions are annotated using a structured, fine-grained protocol, with cross-modal temporal alignment ensured via precise synchronization. Crucially, the dataset employs a gamified, context-aware emotion elicitation paradigm, facilitating future integration of textual (NLP) and other modalities. Contribution/Results: FERG is publicly released, comprising 10 hours of high-fidelity, time-aligned multimodal data. It fills a critical gap in French multimodal emotion resources and significantly enhances model generalizability and robustness in natural human–computer interaction scenarios.
Existing affective computing research predominantly relies on single-modality data, lacking high-synchrony, multi-turn multimodal benchmark datasets for stress response analysis. Method: We introduce the first synchronized multimodal stress dataset, concurrently capturing facial video and nine physiological signals—including heart rate, electrodermal activity, and skin temperature—using a high-frame-rate camera and medical-grade wearable sensors (Empatica E4). Data were collected from 20 participants across 26 hours of ecologically valid stress-inducing scenarios, with rigorous temporal alignment, artifact correction, and multi-source signal-to-noise ratio validation to ensure high fidelity and cross-session consistency. Contribution/Results: A ResNet+LSTM fusion model trained on this dataset achieves 89.3% accuracy in stress-level classification, significantly outperforming unimodal baselines. This work establishes the first high-quality, cross-modal benchmark for stress recognition, addressing a critical gap in affective computing and multimodal behavioral physiology.
This study addresses critical privacy, ethical, and cross-cultural bias challenges arising from integrating affective computing and large language models into AI systems for emotion recognition and response. Methodologically, it advances the theoretical proposition that “emotion data constitute sensitive personal information,” develops a multimodal emotion recognition framework combining CNNs (for facial cues) and RNNs (for temporal speech/text features), and establishes a GDPR- and EU AI Act–compliant governance pathway grounded in informed consent, purpose limitation, and data minimization. Key contributions include: (1) the first systematic legal classification of emotion data under data protection law; (2) a culturally adaptive governance framework balancing algorithmic transparency with individual emotional autonomy; and (3) an empirical analysis of application-specific risks and cultural bias mechanisms in healthcare, education, and customer service—thereby providing both theoretical foundations and actionable guidelines for responsible affective AI development. (149 words)
研究通过构建Mult2EMo数据集,分析社交媒体帖子中的文本和图像如何共同表达情绪,以及读者理解这些情绪的能力,强调了触发事件在情绪理解中的重要性。
This study investigates whether machines can accurately perceive human emotional responses they elicit, thereby testing the affective closed-loop hypothesis. In a virtual reality experiment, participants were exposed to six emotionally charged stimuli derived from critical incidents experienced by emergency responders, while subjective ratings of valence and arousal, along with physiological signals—including electrodermal activity and heart rate—were simultaneously recorded. The findings reveal, for the first time, an “expression–perception asymmetry”: although the machine effectively induced anger, fear, and sadness (dz = 1.1–1.7), self-reported arousal showed no significant change, and physiological measures responded only to the most salient events. Moreover, emotional valence became decoupled from physiological signals, suggesting that current peripheral physiological channels are insufficiently reliable for decoding machine-induced affective states.
研究通过分析观众在戏剧表演中生理信号的变化,使用多模态传感数据和综合分析框架,揭示了不同情感场景下的生理响应模式。
This study addresses the lack of systematic characterization of multimodal affective computing datasets for continuous valence–arousal annotation. We conduct a comprehensive survey of 25 such datasets published between 2008 and 2024, analyzing their scale, participant demographics, sensor modalities (e.g., EEG, ECG, facial video, speech), annotation protocols, and data formats. Through cross-dataset comparative analysis and methodological evaluation, we chart the technical evolution and application distribution of these resources for the first time. Our findings reveal a dominant trend toward camera-centric acquisition coupled with synergistic multimodal fusion, and quantitatively demonstrate the performance gains achievable through integrated physiological–behavioral signal fusion. The study delivers an authoritative, empirically grounded methodology guide for dataset selection, model design, and real-world deployment of affective computing systems—particularly in human–computer interaction, mental health monitoring, and autonomous driving applications.