Score
Quantifying, representing, and exploiting emotional states in interactive systems—e.g., modeling relationships between micro-actions and emotions, representing dynamic affect (valence–arousal) and using those representations to route users to specialized agents.
This paper addresses the core challenge of insufficient machine empathy in affective computing. We propose a unified framework integrating large language models (LLMs), multimodal learning (text, speech, and physiological signals), and personalized modeling. Through a systematic review, we analyze advances in emotion recognition, sentiment analysis, and personality modeling across four key application domains: AI chatbots, multimodal human–computer interaction, mental health interventions, and safety-critical systems—revealing empirical patterns linking data modality, scale, and diversity to model performance. We introduce, for the first time, a comprehensive research paradigm encompassing ethical assessment, annotated dataset analysis, and verifiability-oriented design, thereby clarifying technical trajectories and identifying critical research gaps. Finally, we formulate a tripartite design principle—“safety–empathy–utility”—for next-generation affective support systems, accompanied by an empirically grounded validation pathway.
This study addresses the lack of systematic integration in current Affective Extended Reality (Affective XR) research, which hinders a clear understanding of its design paradigms and technical landscape in emotion recognition and sharing. Employing a scoping review methodology, the work systematically analyzes 82 human-computer interaction studies, synthesizing advancements in biosignal sensing, XR platforms, and affective modeling to produce the first comprehensive research map of Affective XR. The analysis reveals the diversity of emotion-sharing objectives, distills key design dimensions, and categorizes prevailing system architectures and evaluation approaches. Furthermore, it identifies underexplored research directions, offering a cohesive theoretical framework and actionable pathways to guide future investigations in this emerging interdisciplinary domain.
This study addresses the lack of systematic characterization of multimodal affective computing datasets for continuous valence–arousal annotation. We conduct a comprehensive survey of 25 such datasets published between 2008 and 2024, analyzing their scale, participant demographics, sensor modalities (e.g., EEG, ECG, facial video, speech), annotation protocols, and data formats. Through cross-dataset comparative analysis and methodological evaluation, we chart the technical evolution and application distribution of these resources for the first time. Our findings reveal a dominant trend toward camera-centric acquisition coupled with synergistic multimodal fusion, and quantitatively demonstrate the performance gains achievable through integrated physiological–behavioral signal fusion. The study delivers an authoritative, empirically grounded methodology guide for dataset selection, model design, and real-world deployment of affective computing systems—particularly in human–computer interaction, mental health monitoring, and autonomous driving applications.
This study addresses the lack of a standardized approach to visualizing emotional states in existing information systems. Through a systematic user study, it investigates how discrete emotions and dimensional affect models map onto visual variables—including color, size, speed, shape, and animation. The work establishes, for the first time, empirical associations between affective labels and multidimensional visual encodings, revealing that color, size, and speed are significantly correlated with discrete emotions, while speed also exhibits a strong relationship with arousal. These findings provide an empirical foundation for developing universal guidelines for affective visualization and advance the standardization of emotion representation in visual interfaces.
This study addresses a critical limitation of current virtual conversational agents—namely, their rich knowledge base coupled with a notable deficiency in emotional perception and expression, which hinders their ability to adapt to users’ affective states and ultimately constrains interaction quality and user engagement. To overcome this, the work proposes a novel approach that explicitly integrates emotional context into general-purpose dialogue generation by synergistically combining sentiment analysis, natural language processing, and generative AI. The resulting emotion-aware conversational agent not only accurately recognizes user emotions but also produces empathetic and expressive responses. Preliminary empirical evaluations demonstrate that this method significantly enhances user engagement, system usability, and overall dialogue quality compared to conventional content-focused systems.
This study addresses the modeling of emotional states and their short-term dynamics in user-generated text by proposing a unified framework that integrates large language model prompting, an Ising-inspired pairwise maximum entropy transition structure, and lightweight neural regression. The approach introduces trainable user embeddings and temporal affective trajectories to jointly predict the evolution of valence and arousal. Its key innovation lies in coupling a structured emotion transition mechanism with personalized user representations, revealing that affective dynamics are primarily driven by numerical trajectory patterns rather than textual semantics. The proposed system achieved top performance in both Subtask 1 and Subtask 2A of SemEval-2026 Task 2.
This work addresses a critical gap in the evaluation of decision-oriented small language models: the neglect of emotion’s causal influence on behavior. We propose a novel paradigm that integrates emotion induction at the representational level with structured game-theoretic assessment. Our approach enables controlled and transferable emotional interventions through activation manipulation grounded in authentic emotional text. To evaluate model behavior under diverse strategic conditions, we construct a decision-making benchmark encompassing cooperative and competitive settings, as well as complete and incomplete information scenarios, drawing from Diplomacy, StarCraft II, and realistic human personas. Experiments reveal that emotional perturbations systematically alter model strategy selection but often induce behavioral instability or deviations from human expectations. Building on these findings, we further introduce effective methods to enhance the emotional robustness of language models.
This work addresses the modeling challenge of micro-gestures (MGs) in identity-agnostic affective intelligence. It introduces the first systematic definition of MG-based affective semantics, departing from conventional action recognition paradigms. Methodologically, we propose a plug-and-play spatio-temporal balanced fusion module and establish a micro-pose-aware enhancement strategy synergized with large language models for affective reasoning. Our contributions are threefold: (1) We uncover, for the first time, the unique semantic value of MGs in fine-grained, unconscious affective expression; (2) Our approach achieves state-of-the-art performance on MG recognition with strong cross-dataset generalization; (3) It significantly improves the completeness and depth of affective understanding, enabling downstream applications such as deception detection.
Traditional user modeling approaches implicitly handle psychological states, limiting their ability to accurately interpret behavior in long-term, socially interactive settings. This work proposes the Mind Modeling (M3) framework, which for the first time systematically integrates Theory of Mind (ToM) into user modeling by explicitly representing mental states such as beliefs, intentions, emotions, and knowledge. M3 employs an integrated perception–mentalization–action architecture that enables dynamic inference and continuous updating of these states. The approach significantly enhances the interpretability and cross-session consistency of personalized systems. Feasibility is demonstrated through embodied interaction trajectories, establishing M3 as a novel paradigm for next-generation personalization.
This study addresses the lack of systematic understanding regarding affective interaction mechanisms and behavioral stability of large-scale AI agents on social platforms such as Moltbook. The work proposes the first emotion-aware modeling framework for multi-agent interactions, mapping textual exchanges to fine-grained emotional categories through a Persona-Stimulus-Reaction (PSR) paradigm that quantifies the consistency of agents’ emotional responses under similar contextual stimuli. By integrating emotion classification, natural language processing, and context alignment analysis, the research demonstrates that distinct AI agents exhibit unique affective profiles, and their behavioral stability is significantly modulated by interaction context. These findings establish a theoretical foundation for designing trustworthy multi-agent systems with emotionally coherent and contextually adaptive behaviors.
This work addresses the limitation of existing character dialogue systems, which typically treat emotion as a static trait and thus fail to model dynamic emotional evolution triggered by external events. To overcome this, the study introduces the Component Process Model (CPM) from psychology into the field for the first time and proposes CPM-MultiAgent, a multi-agent framework that enables continuous emotional evolution across multi-turn dialogues. The framework integrates emotion trigger recognition, CPM-based collaborative appraisal, and a latent-variable mechanism for updating emotional states. Experimental results demonstrate that the proposed approach significantly enhances both emotional consistency and dynamic expressiveness, as evidenced by automatic metrics, ablation studies, human evaluations, and case analyses, making it particularly suitable for emotion-sensitive applications such as healthcare and education.
Microexpressions, due to their subtle and fleeting nature, have limited human-centered applications in empathy enhancement. This work proposes, for the first time, a microexpression visualization framework grounded in a social augmentation perspective, which leverages computational analysis and real-time visualization techniques to map imperceptible microexpressions into perceivable affective cues, thereby aiming to enrich interpersonal empathic experiences. Rather than focusing solely on microexpression recognition, the framework prioritizes empathy facilitation within human-centered interaction contexts. Although empirical results have not yet been reported, a controlled experimental design has been developed to evaluate its feasibility, offering a novel pathway at the intersection of affective computing and social augmentation research.
This study addresses the challenge of sustaining natural, coherent, and emotionally engaging long-term human–agent interactions, which existing virtual agents struggle to support due to a lack of cross-temporal affective modeling. To bridge this gap, the authors propose the Cross-Temporal Emotion Modeling (CTEM) framework, which establishes a continuous closed loop among behavioral memory, dynamic emotional states, and anticipated future interactions—thereby enabling emotional consistency, reflection, and anticipation. CTEM leverages foundation models to design mechanisms for tracking and updating emotional states, integrating a memory system with user feedback to drive affectively grounded dialogue generation. In a 21-day in-the-wild study, the CTEM-equipped virtual agent Auri significantly enhanced users’ perceived naturalness, coherence, and emotional harmony in interaction.