🤖 AI Summary
This study addresses the challenges of constructing foundation models from unlabeled inertial data for human activity recognition and the absence of global structure in existing approaches. To this end, we propose a morphological contrastive learning framework that pioneers explicit global structure modeling within self-supervised learning for inertial signals. By leveraging motion primitive discovery and domain-specific feature descriptors, our method achieves structure-aware grouping, effectively overcoming the limitations of random sampling while injecting global structural information to optimize pre-training. Experimental results demonstrate that the proposed approach significantly enhances encoder representational capacity and cluster separability in the embedding space, improving F1 scores by up to 15%. Notably, it surpasses existing foundation model performance while requiring only 1/4600th of the labeled data.
📝 Abstract
Despite the ubiquity of sensors in wearable and mobile devices and the abundance of human movement data they generate, translating unlabeled recordings into foundational motion models remains an open challenge. Self-supervised learning (SSL) has alleviated the need for costly annotations, yet existing approaches leave the global structure of large-scale motion data largely untapped, relying on randomly sampled batches and local comparisons that become particularly problematic for in-the-wild inertial data dominated by stationary, low-variance behaviors. Here we introduce Morphological Contrastive Learning (MorphCL), a self-supervised pretraining framework that uses structure-aware grouping to inject explicit modeling of global structure into inertial-based SSL approaches. Building on two well-established pillars of motion analysis, the discovery of motion primitives, or motifs, and domain-specific feature descriptors, we show that MorphCL substantially improves linear probing and finetuning results of learned encoders by up to 15 percentage points in F1-score. In a comparison with existing foundation models, we demonstrate that MorphCL-pretrained encoders match or surpass them models in linear probing performance while trained on $4600\times$ less data. Qualitative analysis of the resulting embedding spaces further reveals morphologically meaningful cluster structure, with improved separation of kinematically similar activity classes.