🤖 AI Summary
This study addresses the challenge of integrating instantaneous emotion inference, regulation, and cross-session continuity in generative agents for sustained emotional support. We propose PAIR, a real-time multimodal agent that innovatively combines perceptive emotion inference with regulation mechanisms. Specifically, PAIR reconstructs user emotional states via an evaluation scaffolding algorithm, delivers personalized guidance through coordinated voice, color, and avatar modalities, and maintains cross-session contextual coherence using rolling memory. A 14-day deployment experiment demonstrates that the system achieves a mean absolute error of 1.20 in valence prediction, significantly alleviates users' negative emotions, and enhances perceived helpfulness. These findings establish PAIR as an effective paradigm for long-term multimodal affective companionship.
📝 Abstract
Sustained emotional support requires generative agents to connect momentary emotion inference and regulation with continuity across encounters. We present PAIR (Perceptual Affective Inference and Regulation), a real-time multimodal agent that reconstructs how an event is appraised into an emotional state. Appraisal scaffolds produce a valence-arousal-dominance estimate and select regulation guidance, delivered through conversation with coordinated speech, color, and avatar cues. Rolling memory carries context across sessions, and the scaffold re-runs after guidance. In a 14-day deployment with 19 participants, 1,093 sessions paired initial and post-guidance estimates with unanchored self-reports. Initial valence reached MAE 1.20 on the 9-point SAM scale (r=.68), dominance reached MAE 1.30, and arousal showed weak agreement even after coarsening. Self-reported emotional change varied with initial state, with the largest valence increases in sessions that began at negative valence. Perceived understanding was associated with greater valence increase and showed little correspondence with numerical prediction error. Over two weeks, helpfulness increased while input shortened; interviews traced personalization and companionship to relevant recall, context updates, and familiar dialogue. These findings connect inference accuracy to conversational and temporal patterns of support through per-event, first-person evaluation.