🤖 AI Summary
This study addresses the limited contribution of physiological signals in EEG-fNIRS-based emotion regression within familiar video contexts. We propose leveraging video temporal priors as a strong baseline model, repositioning physiological signals as optional residual correction terms. The approach is systematically evaluated using leave-subject-out five-fold cross-validation, source-explicit ablation studies, and fixed fusion strategies. Results demonstrate that the low-cost video prior outperforms complex multimodal fusion schemes, achieving mean absolute errors (MAE) of 0.05 and 0.32 for intra- and inter-subject evaluations, respectively. These findings confirm that video identity and temporal information are core factors in continuous emotion prediction, providing an efficient paradigm reference for multimodal affective computing.
📝 Abstract
Continuous emotion regression estimates moment-to-moment valence and arousal while a viewer watches a video. In familiar-video deployment, responses fron training participant-specific estimate, and prior-dominating fixed fusion tests whether physiology adds residual correction. In five-fold subject-held-out evaluation on 24was within 0.05 and 0.32 MAE of fusion in the internal and external evaluations, respectively. Source-explicit ablations showed that video identity and within-video tine accounted for most of the reduction, while EG-FNIRS gains were smaller and varied across participants and videos. These results identify the video-time prior as a strong, low-cost baseline and position EEG-fNIRS as an optional residual signal for familiar-video emotion regression.