Artifact Annotations Partially Substitute for Per-User Calibration: SAFE-EDA and a Normalization-Controlled Evaluation of Wrist-EDA Affect Recognition

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overestimation of pre-training benefits in wrist-worn electrodermal activity (EDA) emotion recognition, where ambiguous normalization statistics obscure the distinction between first-wear and post-calibration performance. We propose SAFE-EDA, a model pre-trained with expert artifact annotations, and systematically evaluate cross-user generalization by controlling normalization sources. Employing a compact convolutional neural network, leave-one-subject-out cross-validation, and multi-dataset comparisons, we reveal the critical impact of normalization strategies on reported outcomes. Results demonstrate that pre-training significantly improves F1 scores when normalization relies exclusively on training-set statistics; however, these gains vanish and become non-significant when individual full-recording statistics are used. Furthermore, this work confirms that artifact-supervised pre-training outperforms self-supervised alternatives, underscoring the necessity of rigorous reporting standards for normalization protocols in physiological computing research.
📝 Abstract
Wrist electrodermal activity (EDA) differs in amplitude from one person to the next, so affect-recognition models normalize their input before classification. Studies that test such models on held-out subjects seldom report where the normalization statistics come from, yet statistics computed from the held-out subject's own recording give the model information that a device does not have when it is first worn. We asked how this choice alters the measured benefit of pretraining. A compact convolutional network, SAFE-EDA, was pretrained on expert artifact annotations from 43 subjects and compared with the same network trained from scratch on the Wearable Stress and Affect Detection (WESAD) dataset (15 subjects, leave-one-subject-out), with two normalization sources crossed with four window hops. When the statistics came only from training subjects, pretraining raised macro-F1 by 0.078 to 0.227; when they came from the held-out user's full recording, the gain fell to between 0.020 and 0.050 and was no longer significant. Artifact supervision was far more useful than self-supervised pretraining on the same recordings (0.078 versus 0.008). Across 13 configurations in two datasets, the pretrained network was better in 12, but on the second dataset (26 subjects) per-user normalization increased the gain instead of reducing it, so the interaction depends on the data. Only five of 50 published WESAD studies state which data were used for normalization. Reporting this choice is necessary to separate first-use performance from performance after calibration.
Problem

Research questions and friction points this paper is trying to address.

Electrodermal Activity
Affect Recognition
Normalization
Pretraining Evaluation
Per-User Calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Electrodermal Activity
Affect Recognition
Artifact Annotations
Normalization-controlled Evaluation
Pretraining
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haochen Chai
College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China
X
Xinbi Luo
College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China
Zining Liu
Zining Liu
University of Pennsylvania
F
Fangfang Jiang
College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China