🤖 AI Summary
This work addresses the limitations of existing IMU-based perception methods, which rely on costly labeled data and exhibit poor generalization across devices, sensor placements, and users. The authors propose a physics-informed self-supervised learning framework that replaces conventional neural decoders with an adaptive physics decoder, enabling two-stage encoding and reconstruction within a structured latent space while explicitly incorporating physical priors without any labels. Key innovations include probabilistic frequency–spatial constraints to disentangle motion components, multi-view kinematic tree modeling, and an uncertainty-aware mechanism. Evaluated on inertial tracking and full-body motion capture tasks, the method reduces error by up to 5× and 4×, respectively, under challenging cross-domain generalization scenarios, significantly outperforming current supervised and self-supervised baselines.
📝 Abstract
Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning, an autoencoder-style paradigm for label-free IMU sensing. We replace the conventional neural decoder with an auto-adaptive physics decoder, a learnable family of kinematic equations that enforces explicit physical structure while adapting across environments, and adopt a hybrid two-stage IMU encoder with reconstruction in a structured latent space to mitigate sensor noise. Our framework further introduces probabilistic frequency-spatial constraints to disentangle sensor and object motion, a multi-view kinematic tree to exploit sparse physical self-supervised signals, and an uncertainty-aware formulation to handle the inherent ambiguity of IMU inference. Evaluated on inertial tracking and full-body motion capture over public datasets and realistic deployments, physical self-supervised learning reduces errors by up to 5x for tracking and 4x for motion capture in challenging generalization scenarios, consistently outperforming state-of-the-art supervised and self-supervised baselines without any labels.