Score
Designs and implements models and architectures for predicting, forecasting, or generatively modeling sequences with explicit time awareness, including temporal interval embeddings and mechanisms for irregularly sampled, long-range, hierarchical, or multi-turn sequences and for incorporating multimodal inputs. Builds probabilistic, conditional, masked, and neural sequence models (e.g., time-aware transformers, GRUs, hierarchical and long-sequence architectures) for sequence conditioning, uncertainty-aware forecasting, and generation over variable-length temporal data.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
Transformers suffer from inadequate sequential representation learning in time-series forecasting, struggling to model multivariate channel correlations and handle data instability. Method: This paper proposes Learnable Sequence Complementors—a novel mechanism that augments input sequences to enhance representational diversity. Grounded in an information-theoretic analysis, it first establishes a linear relationship between sequence diversity (measured by entropy) and prediction error. Leveraging this insight, we design a provably sound complement attention mechanism and introduce a theoretically grounded diversity loss, jointly optimized with entropy regularization. Contribution/Results: Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art methods on both long- and short-term forecasting benchmarks, achieving significant reductions in MSE. The results empirically validate that increasing sequential representation diversity is critical for improving forecasting accuracy.
Structured State Space Models (SSMs) face a fundamental trade-off between long-range dependency modeling and computational efficiency, limiting their broad adoption across NLP, speech, vision, and time-series domains. This paper provides the first systematic survey of SSMs—from theoretical foundations (continuous-time dynamics, HiPPO projections) to industrial variants (S4, Mamba, S5, Jamba)—unifying analysis of their linear-time complexity, memory-efficient parameterization, and hardware-aware inference acceleration. We identify selectivity mechanisms and low-rank structured matrices as key innovations enabling SSMs to emerge as the third major sequence modeling paradigm—alongside RNNs and Transformers—achieving near-Transformer accuracy on long-sequence tasks while reducing memory footprint by over 70% and significantly improving inference throughput. We further highlight critical open challenges: training instability, hybrid modeling strategies, and interpretability.
This work addresses the lack of systematic evaluation of sequence models’ ability to capture diverse temporal dependencies—such as short- and long-range, decaying, and oscillatory patterns. We propose the first synthetic benchmark framework based on controllable, parameterized memory functions. By explicitly designing memory kernel functions, our framework generates synthetic tasks with continuous-time complexity, enabling fine-grained, interpretable, and theoretically grounded analysis of model memory characteristics. We evaluate mainstream architectures—including RNNs, Transformers, and State Space Models (SSMs)—under a unified benchmark across multiple dimensions. Our experiments not only validate existing theoretical predictions but also uncover, for the first time, implicit architectural preferences for specific memory patterns and their precise failure boundaries. The results provide reproducible, interpretable, and quantitative guidance for selecting appropriate sequence modeling paradigms.
Despite theoretical support for infinite context, state space models (SSMs), linear RNNs, and other sequence architectures exhibit substantially degraded performance on ultra-long sequences in practice, with large inter-architectural disparities in extrapolation capability. Method: We conduct the first systematic empirical evaluation of SSMs, linear RNNs, and Transformer variants across controlled synthetic tasks and real-world long-text benchmarks, analyzing their context scaling behavior and generalization curves. Results: All models suffer sharp performance drops beyond certain sequence lengths, indicating a fundamental gap between theoretical infinite-context capacity and empirical efficacy. Crucially, inductive bias—not parameter count or training scale—emerges as the dominant factor governing practical long-range modeling effectiveness. Our findings challenge prevailing assumptions about asymptotic context scalability and provide attributable, evidence-based insights into the failure mechanisms of long-range dependency modeling.
This work proposes a probabilistic inference framework that integrates inductive biases to address the challenges of uncertainty quantification in deep sequential models. While traditional Bayesian approaches struggle with prior specification and inference accuracy in large-scale networks, the proposed method establishes a theoretical connection between Transformer attention mechanisms and sparse Gaussian processes, enabling scalable approximate Bayesian inference. It introduces cross-domain inducing points derived from HiPPO operators to support long-range historical modeling in online learning settings. Furthermore, self-supervised signals are leveraged to enrich the probabilistic structure of latent variables in sequence generation. The resulting approach significantly enhances the uncertainty quantification capability, probabilistic expressiveness, and scalability of deep sequential models, all while maintaining competitive predictive performance.
Existing sequence models often neglect explicit temporal information, limiting their capacity for temporal reasoning and event timing modeling. This work proposes ChronoSSM, an autoregressive state space model that jointly models event content and timestamps within the SSM framework for the first time. By employing a shared backbone network, ChronoSSM simultaneously optimizes event prediction and time generation objectives, departing from conventional two-stage paradigms and enabling learned representations to encapsulate both semantic and temporal structures. Experiments across four datasets with varying levels of temporal annotation density demonstrate that ChronoSSM substantially improves the recoverability of inter-event timing information while maintaining strong event generation quality.
This work addresses the limitations of conventional Transformers in time series modeling—namely, their lack of interpretability, inadequate channel-wise modeling capacity, and absence of structured conditional generation mechanisms. The authors propose the first systematic extension of Probabilistic Transformers (PT) to Spatio-Temporal Probabilistic Transformers (ST-PT), formulated as programmable factor graphs. By explicitly designing graph topology, conditional potential functions, and a message-passing mechanism driven by variational inference, ST-PT achieves three key innovations: incorporating structural priors to enhance few-shot performance, enabling external conditions to structurally govern the generative process, and integrating CRF-based teacher distillation of latent variables to mitigate autoregressive error accumulation. Experimental results demonstrate that ST-PT serves as a flexible and effective general-purpose backbone for spatio-temporal modeling.
Existing LLM-based time series forecasting methods overlook the statistical properties and dynamic dependencies inherent in temporal data. To address this, we propose MAP4TS—a multi-prompt fusion framework tailored for time series analysis—that explicitly encodes autocorrelation (ACF), partial autocorrelation (PACF), and Fourier spectral features as statistical prompts, and synergistically integrates them with global/local domain prompts and temporal structural prompts to bridge classical time series analysis with LLM-based reasoning. A cross-modal alignment module fuses handcrafted statistical features with raw time series embeddings. Evaluated on eight benchmark datasets, MAP4TS significantly outperforms state-of-the-art LLM-based forecasters; notably, even a lightweight GPT-2 achieves superior long-horizon forecasting accuracy compared to large LLMs. Ablation studies confirm the critical and complementary roles of all four prompt components in enhancing both prediction stability and accuracy.
This work proposes an efficient hierarchical sequence modeling architecture based on a binary tree structure to address the high computational complexity of standard self-attention and its difficulty in capturing hierarchical dependencies in long sequences. The method introduces a binary tree reduction mechanism with hierarchical inductive bias, replacing conventional self-attention with recursively applied Gated Linear Units (GLUs). This design achieves O(log n) parallel depth while maintaining O(n) space complexity. Experimental results demonstrate that the proposed model significantly outperforms standard Transformers on long-sequence tasks, exhibiting faster convergence and higher accuracy—particularly excelling in tasks where hierarchical dependency structures are critical.