🤖 AI Summary
This work addresses the limitations of conventional convolutional time series models in effectively capturing multi-scale structures within long sequences and the tendency of pooling operations to discard critical temporal positional information. The authors propose the ROMAN operator, which explicitly encodes temporal scale and coarse-grained time positions into the channel dimension for the first time. By integrating anti-aliased multi-scale pyramid decomposition, fixed-window slicing, and channel stacking, ROMAN significantly reduces sequence length while preserving multi-scale interactions and temporal positional cues. This design introduces a controllable inductive bias into downstream standard convolutional classifiers such as MiniRocket and FCN. Experiments demonstrate that ROMAN successfully captures coarse-grained positional and multi-scale dependencies on synthetic tasks and substantially improves computational efficiency on long-sequence UCR/UEA benchmarks, with accuracy gains varying across tasks.
📝 Abstract
We introduce ROMAN (ROuting Multiscale representAtioN), a deterministic operator for time series that maps temporal scale and coarse temporal position into an explicit channel structure while reducing sequence length. ROMAN builds an anti-aliased multiscale pyramid, extracts fixed-length windows from each scale, and stacks them as pseudochannels, yielding a compact representation on which standard convolutional classifiers can operate. In this way, ROMAN provides a simple mechanism to control the inductive bias of downstream models: it can reduce temporal invariance, make temporal pooling implicitly coarse-position-aware, and expose multiscale interactions through channel mixing, while often improving computational efficiency by shortening the processed time axis. We formally analyze the ROMAN operator and then evaluate it in two complementary ways by measuring its impact as a preprocessing step for four representative convolutional classifiers: MiniRocket, MultiRocket, a standard CNN-based classifier, and a fully convolutional network (FCN) classifier. First, we design synthetic time series classification tasks that isolate coarse position awareness, long-range correlation, multiscale interaction, and full positional invariance, showing that ROMAN behaves consistently with its intended mechanism and is most useful when class information depends on temporal structure that standard pooled convolution tends to suppress. Second, we benchmark the same models with and without ROMAN on long-sequence subsets of the UCR and UEA archives, showing that ROMAN provides a practically useful alternative representation whose effect on accuracy is task-dependent, but whose effect on efficiency is often favorable. Code is available at https://github.com/gon-uri/ROMAN