🤖 AI Summary
To address the limited local sensitivity and lack of inductive bias in Transformer-based inertial odometry models—which constrain localization accuracy and generalization—this paper proposes IONext, a novel CNN architecture. Its key contributions are: (1) a Dual-branch Adaptive Dynamic Mixing (DADM) module that enables input-driven multi-scale feature aggregation; and (2) a Spatio-Temporal Gating Unit (STGU) integrating large-kernel convolutions, Transformer-like structures, dynamic weight generation, and spatio-temporally decoupled gating, thereby jointly capturing fine-grained local patterns and long-range motion dynamics. Evaluated on six public benchmarks, IONext achieves state-of-the-art performance across all metrics. Notably, on the RNIN dataset, it reduces the average absolute translation error (ATE) and relative translation error (RTE) by 10% and 12%, respectively.
📝 Abstract
Researchers have increasingly adopted Transformer-based models for inertial odometry. While Transformers excel at modeling long-range dependencies, their limited sensitivity to local, fine-grained motion variations and lack of inherent inductive biases often hinder localization accuracy and generalization. Recent studies have shown that incorporating large-kernel convolutions and Transformer-inspired architectural designs into CNN can effectively expand the receptive field, thereby improving global motion perception. Motivated by these insights, we propose a novel CNN-based module called the Dual-wing Adaptive Dynamic Mixer (DADM), which adaptively captures both global motion patterns and local, fine-grained motion features from dynamic inputs. This module dynamically generates selective weights based on the input, enabling efficient multi-scale feature aggregation. To further improve temporal modeling, we introduce the Spatio-Temporal Gating Unit (STGU), which selectively extracts representative and task-relevant motion features in the temporal domain. This unit addresses the limitations of temporal modeling observed in existing CNN approaches. Built upon DADM and STGU, we present a new CNN-based inertial odometry backbone, named Next Era of Inertial Odometry (IONext). Extensive experiments on six public datasets demonstrate that IONext consistently outperforms state-of-the-art (SOTA) Transformer- and CNN-based methods. For instance, on the RNIN dataset, IONext reduces the average ATE by 10% and the average RTE by 12% compared to the representative model iMOT.