🤖 AI Summary
This work addresses two key challenges in multimodal motion recognition: cross-dataset sensor format heterogeneity and time-consuming hyperparameter optimization. We propose the first end-to-end, fully automated, general-purpose framework for this task. The framework uniformly processes heterogeneous time-series sensor data, enabling fully automated preprocessing—including multimodal temporal alignment—model training, Optuna-based hyperparameter optimization, and standardized evaluation—without manual intervention. Its core innovations are zero-configuration cross-dataset transferability and plug-and-play edge deployment. Evaluated on ten diverse, heterogeneous datasets, the framework achieves state-of-the-art performance, improving average classification accuracy by 3.2% and accelerating end-to-end deployment throughput by 8× compared to conventional pipelines.
📝 Abstract
In this paper, we present an end-to-end automated motion recognition (AutoMR) pipeline designed for multimodal datasets. The proposed framework seamlessly integrates data preprocessing, model training, hyperparameter tuning, and evaluation, enabling robust performance across diverse scenarios. Our approach addresses two primary challenges: 1) variability in sensor data formats and parameters across datasets, which traditionally requires task-specific machine learning implementations, and 2) the complexity and time consumption of hyperparameter tuning for optimal model performance. Our library features an all-in-one solution incorporating QuartzNet as the core model, automated hyperparameter tuning, and comprehensive metrics tracking. Extensive experiments demonstrate its effectiveness on 10 diverse datasets, achieving state-of-the-art performance. This work lays a solid foundation for deploying motion-capture solutions across varied real-world applications.