π€ AI Summary
Traditional PPG-based sleep staging significantly underperforms EEG-based methods due to the mismatch between fixed 30-second epoch labels and the high temporal resolution of physiological signals. To address this limitation, this work proposes a hidden semi-Markov modelβbased label expansion mechanism that refines coarse-grained sleep stage annotations to a second-level granularity, explicitly modeling rapid physiological transitions at stage boundaries. By integrating heart rate variability and pulse morphological features, the approach improves staging accuracy by 3.7β5.7 percentage points across four deep learning architectures on the MESA dataset. Furthermore, it demonstrates strong zero-shot transfer generalization on the CFS dataset, effectively overcoming the constraints imposed by conventional fixed-window supervision.
π Abstract
Automated sleep staging assigns discrete stage labels to successive time epochs throughout an overnight recording; conventionally each window spans at least 30 seconds, reflecting the minimum temporal resolution of the clinical scoring standard. Wearable photoplethysmography (PPG) has attracted sustained interest as an ambulatory alternative to laboratory-based polysomnography, which relies on electroencephalography (EEG) and other recording modalities that are impractical outside clinical environments. Yet PPG-based staging trails EEG-based methods by a substantial margin, and we argue this gap largely reflects a mismatch between signal and task. Within a stable stage, PPG's inter-stage feature differences are more subtle than those in EEG; yet at stage boundaries, PPG's principal cardiovascular features, heart rate variability and pulse morphology, shift sharply within seconds. The conventional practice of assigning one label to each 30-second epoch therefore suppresses feature that is concentrated near boundaries. We address this gap in two steps. First, we develop a label expansion pipeline based on Hidden Semi-Markov Models that converts coarse epoch labels into sec-level annotations. To assess whether these expanded labels are reliable enough for downstream supervision, we validate them on a separate expert-reviewed dataset and through an auxiliary sleep-wake task whose labels are independent of the expansion pipeline. Second, we use the resulting sec-level supervision on MESA to improve conventional four-class epoch-level staging across four architecturally diverse baselines by 3.7--5.7\,pp in accuracy against the original epoch labels, with supplementary zero-shot evaluation on CFS showing that the transfer benefit persists under cohort and annotation-protocol shift.