Progressive Memory Transformer: Memory-Aware Attention for Time-Series

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing methods fail to explicitly exploit the multi-scale hierarchical structure of time series by proposing the Progressive Memory Transformer (PMT) framework. PMT introduces a novel explicit multi-scale hierarchical supervision mechanism that reinforces structural representations through independent local, mid-range, and global objectives. Additionally, it designs a window-aligned writable memory module to effectively compensate for the inherent shortcomings of Transformers in mid-range modeling. By integrating multi-scale contrastive losses with an attention-based self-supervised learning strategy, PMT achieves superior performance across multiple classification and forecasting benchmarks. Notably, it demonstrates robust classification capabilities even in low-label regimes, thereby validating its effectiveness in capturing mid-range temporal patterns.
📝 Abstract
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5\% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.
Problem

Research questions and friction points this paper is trying to address.

Time-series
Self-supervised learning
Multi-scale representation
Structural hierarchy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progressive Memory Transformer
Multi-scale Representation
Self-supervised Learning
Memory-Aware Attention
Time-Series
💼 Related Jobs
No related jobs found.
T
Tord Sture Stangeland
University of Oslo, Department of Informatics, Oslo, Norway; NORSAR, Lillestrøm, Norway
A
Andreas Köhler
NORSAR, Lillestrøm, Norway; University of Tromsø, Department of Geosciences, Tromsø, Norway
S
Steffen Mæland
NORSAR, Lillestrøm, Norway; Western Norway University of Applied Sciences, Bergen, Norway
Adín Ramírez Rivera
Adín Ramírez Rivera
Professor, University of Oslo
Image ProcessingComputer VisionMachine Learning