🤖 AI Summary
To address the dual challenges of resource constraints in IoT devices and insufficient long-term temporal modeling in inertial navigation, this paper proposes a lightweight yet high-accuracy neural network framework. The method employs a multi-branch reparameterized architecture to enhance feature representation, which collapses into a single-path structure during inference to minimize computational overhead. It further introduces a time-scale sparse attention mechanism to effectively capture long-range temporal dependencies in motion trajectories, and integrates gated convolutional units to jointly fuse local spatial details with distant contextual information. Evaluated on the RoNIN dataset, the proposed approach achieves a 2.59% absolute reduction in trajectory error compared to ResNet, while reducing parameter count by 3.86%, thereby significantly improving the accuracy–efficiency trade-off. The core contributions lie in the synergistic design of (i) a reparameterized single-path inference architecture, (ii) sparse time-scale attention, and (iii) gated convolutional fusion.
📝 Abstract
Inertial localization is regarded as a promising positioning solution for consumer-grade IoT devices due to its cost-effectiveness and independence from external infrastructure. However, data-driven inertial localization methods often rely on increasingly complex network architectures to improve accuracy, which challenges the limited computational resources of IoT devices. Moreover, these methods frequently overlook the importance of modeling long-term dependencies in inertial measurements - a critical factor for accurate trajectory reconstruction - thereby limiting localization performance. To address these challenges, we propose a reparameterized inertial localization network that uses a multi-branch structure during training to enhance feature extraction. At inference time, this structure is transformed into an equivalent single-path architecture to improve parameter efficiency. To further capture long-term dependencies in motion trajectories, we introduce a temporal-scale sparse attention mechanism that selectively emphasizes key trajectory segments while suppressing noise. Additionally, a gated convolutional unit is incorporated to effectively integrate long-range dependencies with local fine-grained features. Extensive experiments on public benchmarks demonstrate that our method achieves a favorable trade-off between accuracy and model compactness. For example, on the RoNIN dataset, our approach reduces the Absolute Trajectory Error (ATE) by 2.59% compared to RoNIN-ResNet while reducing the number of parameters by 3.86%.