🤖 AI Summary
This work addresses the challenge of simultaneously achieving task classification, disentangled representation learning, multimodal fusion, and generative modeling in multimodal wearable computing. To this end, we propose OmniDecVAEs, a lightweight, modality-agnostic unified framework that leverages shared asymmetric autoencoders and a multi-view self-supervised disentanglement loss to learn modality-conditioned time-frequency latent subspaces. With a fixed parameter budget, the method enables high-accuracy recognition, high-fidelity generation, and real-time inference. Evaluated on a 30-modality human activity recognition benchmark, our approach improves activity and subject identification accuracy by 1.01% and 6.75%, respectively, reduces reconstruction error by 76.84%, and enhances synthetic data distribution similarity by 13.85%.
📝 Abstract
Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they do not simultaneously address task-specific classification performance, disentangled and interpretable representation learning, fusion, and generative modeling of highly heterogeneous multi-modal time series. To address this gap, we introduce Omni-modal Variational Decomposition Autoencoders (OmniDecVAEs), a framework that efficiently learns multi-purpose representations in a unified and scalable manner from arbitrarily many modalities. OmniDecVAEs extend DecVAEs by learning modality-conditioned time-frequency latent subspaces through a multi-view self-supervised decomposition loss and a shared asymmetric autoencoder (AE) architecture. Results on a challenging omni-modal human activity recognition (HAR) setting with up to thirty modalities, demonstrate the ability of OmniDecVAEs to learn full-stack wearable representations. When compared to transformer-based and VAE-based methods, OmniDecVAEs full-stack disentangled representation properties lead to accuracy improvements of 1.01% and 6.75% in activity and identity recognition, respectively. Furthermore, OmniDecVAEs synthesize realistic omni-modal time-frequency data that manifest with enhanced reconstructions (mean absolute error improves by 76.84%) and distributional similarity between real and synthetic data (maximum mean discrepancy improves by 13.85%). Our results highlight OmniDecVAEs potential as a lightweight model suitable for intelligent edge wearables and clinical healthcare, unifying processing requirements and abilities in a single model, through its enhanced representational capacity, modality-invariant spatial complexity (4.1M parameters), and real-time latency.