🤖 AI Summary
This study addresses the scarcity of annotations in contactless cardiopulmonary sensing and the limited applicability of contrastive learning to physiological signals by proposing a phenomenon graph-based Joint Embedding Predictive Architecture (JEPA). The method integrates multimodal radar and visual data for self-supervised pretraining, eliminating the need for negative samples and artificial augmentations. By leveraging temporal convolutions, frequency-domain encoding, and Takens' delay embedding, it utilizes the structural properties of physiological phenomenon graphs to guide predictive tasks while validating the effectiveness boundaries of prior assumptions. Evaluated on the OMuSense dataset, the proposed model significantly enhances label efficiency, outperforming supervised baselines by 3.91 percentage points.
📝 Abstract
Millimeter-wave (mmWave) radar and RGB-D cameras can record cardiac and respiratory waveforms continuously and without contact, but labeled recordings remain scarce because every label requires a supervised acquisition session. Self-supervised pretraining can exploit the unlabeled signals, yet contrastive methods depend on signal transformations and negative pairs whose validity is uncertain for cardiorespiratory data, where time warping changes breathing rate and distant windows can share the same physiological state. We present Phenomenon-Graph JEPA, a joint-embedding predictive architecture that learns from four processed one-dimensional streams without negative pairs or synthetic augmentation in its base configuration. Each stream is encoded by a temporal convolutional branch and a band-limited spectral branch. During pretraining, the model predicts stopped target embeddings along typed edges, which connect streams assigned to the same physiological phenomenon, and forward in time within a state episode. We treat this physiological typing as a testable hypothesis and compare it with wrong-edge and all-pairs prediction graphs. In the OMuSense-23 dataset, pretraining improves label-efficiency area over matched supervised training by 3.91 percentage points (95% interval 2.08 to 5.80, Holm-adjusted p = 0.006), and by 3.74 points under a second configuration evaluated on the same test participants. However, the wrong-edge and all-pairs controls do not establish a benefit from physiological typing. Optional Takens-inspired delay coordinates improve a validation comparison with learned history, whereas two wrist-only WESAD protocols do not establish a pretraining advantage. The study therefore separates the measured benefit of predictive representations from the physiological prior used to organize their training.