🤖 AI Summary
This study addresses the limitation of existing photoplethysmography (PPG) foundation models in neglecting multi-duration feature discrepancies for self-supervised learning. We propose a novel "retain-expand" pretraining paradigm built upon a Transformer architecture. By progressively extending supervisory signals across durations ranging from 10 to 240 seconds, and leveraging parameter group freezing and reuse mechanisms, a single encoder incrementally acquires multi-scale representations while effectively mitigating catastrophic forgetting. Experimental evaluations demonstrate that the proposed model outperforms five mainstream PPG foundation models on 12 of 18 tasks across eight datasets. These results validate the effectiveness and superiority of the multi-granularity continual pretraining strategy for developing robust PPG foundation models.
📝 Abstract
Signal features derived from photoplethysmography (PPG) require different signal durations to characterize. Existing PPG foundation models treat duration as a pretraining or evaluation condition rather than using the different durations required by PPG features to organize self-supervision. We hypothesize that self-supervision should expand with signal duration, allowing a single encoder to progressively acquire additional features while preserving and reusing earlier learning. We introduce Retain-and-Extend PPG (RAE-PPG), which trains a single Transformer encoder successively on 10 s, 30 s, and 240 s inputs, adding supervision for signal features supported by each longer observation. The encoder is partitioned into duration-specific parameter groups, allowing later stages to reuse earlier groups while updating only the group assigned to the current stage. Selected earlier targets are reused to supervise later stages, encouraging the corresponding features to remain accessible in longer-input representations. Direct decoding from the final encoder shows that earlier features remain recoverable from longer-input representations, while later-stage features show higher mean decoding performance at their introduction durations. Controlled comparisons further show that prior-stage learning provides a better basis for learning newly introduced features at both transitions. Across 18 tasks from eight datasets, the final frozen encoder achieves the best observed score on 12 tasks compared with five existing PPG foundation models.