๐ค AI Summary
This study addresses the challenge of inter-subject variability in zero-shot cross-subject continuous emotion regression, aiming to achieve precise valence and arousal prediction for unlabeled individuals using EEG and fNIRS signals. To this end, we propose an individualized calibration framework based on unsupervised physiological markers. Specifically, structured modeling is employed to disentangle population-shared features from subject-specific representations, while alpha-band cross-channel synchrony serves as a physiological anchor to calibrate individual trajectories. This mechanism significantly outperforms conventional end-to-end models, validating the efficacy of single-marker calibration. Experimental results demonstrate that the mean absolute error (MAE) is reduced to 25.96 and 22.80, substantially surpassing baseline methods. Overall, this work provides an efficient new paradigm for cross-subject multimodal affective computing.
๐ Abstract
Continuous, second-by-second valence-arousal estimation from physiological signals is typically studied in a subject-dependent setting, where the model sees labeled data from the same person it is later evaluated on. We study the harder zero-shot cross-subject variant on a synchronized EEG-fNIRS dataset: predict raw-scale ([1, 255]) valence and arousal trajectories for subjects whose labels the model never observes, given only their unlabeled EEG/fNIRS recordings while watching the same video stimuli as a disjoint set of training subjects. We decompose the affect trajectory into a structure shared across subjects who watch the same stimuli and an individual structure estimated for each test subject from a label-free EEG marker (alpha-band cross-channel synchrony), which rescales the shared trajectory around the scale midpoint. We validate the per-subject calibration mechanism on four independent axes: leave-one-subject-out correlation between the marker and each subject's true optimal gain, a functional-form comparison against non-linear alternatives, a repeated leave-4-out component ablation isolating each part of the pipeline's contribution, and a ceiling analysis bounding the remaining headroom for per-subject scaling. On held-out subjects, the model reaches an overall MAE of 25.96 / 22.80 across two evaluation batches (valence 21.94 / 19.6, arousal 29.98 / 26.0), well below EEGNet and ASAC-Net baselines reported for the same subject-independent split (raw scale score 60.6 and 55.0 respectively). We further report a systematic negative-result search across model architectures, feature representations, and prediction targets that found no signal able to improve on the single alpha-synchrony marker.