🤖 AI Summary
Modeling long-range temporal dependencies in high-dimensional sequential data (e.g., video frames) remains challenging for memory-intensive tasks. Method: This paper proposes a memory architecture integrating temporal kernels with dense Hopfield networks. It explicitly encodes temporal bias via a learnable temporal kernel function $K(m,k)$ and establishes a higher-order interaction mechanism grounded in energy-functional optimization, theoretically enabling exponential memory capacity and supporting efficient, continuous retrieval over long sequences. Crucially, it augments the self-attention mechanism—unlike standard Transformers—with a trainable, temporally aware memory module. Contribution/Results: Experiments on movie frame storage and ordered reconstruction demonstrate significant performance gains over baselines, effectively alleviating the long-range dependency bottleneck. The approach offers a novel paradigm for video understanding and long-context modeling in sequence learning.
📝 Abstract
In this study we introduce a novel energy functional for long-sequence memory, building upon the framework of dense Hopfield networks which achieves exponential storage capacity through higher-order interactions. Building upon earlier work on long-sequence Hopfield memory models, we propose a temporal kernal $K(m, k)$ to incorporate temporal dependencies, enabling efficient sequential retrieval of patterns over extended sequences. We demonstrate the successful application of this technique for the storage and sequential retrieval of movies frames which are well suited for this because of the high dimensional vectors that make up each frame creating enough variation between even sequential frames in the high dimensional space. The technique has applications in modern transformer architectures, including efficient long-sequence modeling, memory augmentation, improved attention with temporal bias, and enhanced handling of long-term dependencies in time-series data. Our model offers a promising approach to address the limitations of transformers in long-context tasks, with potential implications for natural language processing, forecasting, and beyond.