Score
Designs and evaluates representation schemes or model components that encode temporal data by combining chronological (time-ordered) and state-ordered (rhythm- or decoded-state-based) scans. These encodings capture local temporal continuity while also representing long-range, rhythm-consistent dependencies by integrating decoded rhythm/state information into the sequence representation.
Structured State Space Models (SSMs) face a fundamental trade-off between long-range dependency modeling and computational efficiency, limiting their broad adoption across NLP, speech, vision, and time-series domains. This paper provides the first systematic survey of SSMs—from theoretical foundations (continuous-time dynamics, HiPPO projections) to industrial variants (S4, Mamba, S5, Jamba)—unifying analysis of their linear-time complexity, memory-efficient parameterization, and hardware-aware inference acceleration. We identify selectivity mechanisms and low-rank structured matrices as key innovations enabling SSMs to emerge as the third major sequence modeling paradigm—alongside RNNs and Transformers—achieving near-Transformer accuracy on long-sequence tasks while reducing memory footprint by over 70% and significantly improving inference throughput. We further highlight critical open challenges: training instability, hybrid modeling strategies, and interpretability.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
This study systematically investigates the necessity and effectiveness of temporal modeling in sketch representation learning: Is it essential to treat sketches as sequences, and which types of sequential information—such as coordinate encoding (absolute vs. relative), decoding paradigm (autoregressive vs. non-autoregressive), or positional encoding—most critically impact representation quality? Through controlled ablation experiments, we evaluate these factors across multiple tasks—including classification, retrieval, and generation—on standard multi-task benchmarks. Results show that: (1) absolute coordinates yield significantly better performance than relative coordinates; (2) non-autoregressive decoders consistently outperform autoregressive ones across all evaluated tasks; and (3) the benefit of temporal modeling is highly task- and sequence-dependent, not universally advantageous. To our knowledge, this is the first work to empirically demonstrate the *conditional validity* of temporal modeling for sketches, providing reproducible, lightweight design principles for efficient sketch representation learning.
This study systematically investigates the frequency-domain encoding capabilities of the Chronos foundation model, addressing a critical gap in understanding how such models represent fundamental signal properties. Through controlled experiments using discrete sinusoidal signals and a lightweight online Minimum Description Length (MDL) probing framework, the work examines the existence, separability, and cross-spectral fidelity of internal frequency representations within the Chronos decoder. The research reveals, for the first time, a degradation in representation quality in high-frequency regions, thereby delineating both the strengths and limitations of Chronos’s frequency encoding mechanism. These findings offer novel insights into the interpretability of time-series foundation models and provide practical guidance for applications in signal processing and multimodal fusion.
This study addresses the limitations of existing self-supervised electrocardiogram (ECG) models, which often rely on short segments and discretized tokens, thereby discarding long-range temporal information critical for rhythm inference. The authors systematically evaluate the impact of varying temporal context lengths—from 16 seconds to 10 minutes—and front-end encoding strategies—continuous convolutional embeddings versus vector-quantized tokens—on representation learning, while keeping the Transformer backbone and training protocol fixed. Model performance is assessed through arrhythmia detection and patient-level retrieval tasks on the Icentia11k dataset. Results demonstrate that extended context significantly enhances transferability and cross-session stability, with models using 5–10 minutes of data achieving optimal performance. Moreover, continuous embeddings consistently outperform discrete tokens across all scales, yielding substantial gains in rhythm detection accuracy and retrieval precision, thereby offering crucial insights for designing clinically oriented ECG foundation models.
This work investigates how time-series foundation models (TSFMs) internally represent fundamental temporal concepts. Addressing four key problems—hierarchical concept encoding, linear recoverability of atomic concepts, depth-wise evolution of representations, and compositional concept handling—we propose a systematic probing framework integrating layer-wise linear probes, representation similarity analysis, and concept disentanglement metrics. We find that TSFMs exhibit a clear semantic hierarchy: shallow layers capture local temporal patterns, while deeper layers encode discrete events and change-point signals. Atomic concepts are reliably localized, yet spectral and deformation factors prove most challenging to recover linearly. Probe performance degrades significantly for composite concepts, revealing representational interference as a bottleneck in modeling interactive dynamics. This study provides the first empirical characterization of temporal semantic hierarchies in TSFMs, offering theoretical foundations for interpretable modeling and architectural refinement.
The internal representation mechanisms of Time Series Foundation Models (TSFMs) remain poorly understood, with limited systematic investigation into inter-layer redundancy and the distribution of interpretable temporal concepts. Method: We first identify a prevalent block-wise redundancy structure across hidden layers; propose an analytical framework grounded in representation similarity and inter-layer self-similarity metrics; construct disentangled representations of temporal concepts—such as periodicity and trend—in latent space; and design concept-level latent-space steering and structured pruning strategies. Contribution/Results: Experiments show that our pruning strategy reduces inference latency by 32%; moreover, we demonstrate precise injection of periodic or trend features into signals originally lacking them, validating both the efficacy and generalizability of concept-level manipulation. This work establishes a novel paradigm for interpretable modeling and controllable editing of TSFMs.
Existing sequence models often neglect explicit temporal information, limiting their capacity for temporal reasoning and event timing modeling. This work proposes ChronoSSM, an autoregressive state space model that jointly models event content and timestamps within the SSM framework for the first time. By employing a shared backbone network, ChronoSSM simultaneously optimizes event prediction and time generation objectives, departing from conventional two-stage paradigms and enabling learned representations to encapsulate both semantic and temporal structures. Experiments across four datasets with varying levels of temporal annotation density demonstrate that ChronoSSM substantially improves the recoverability of inter-event timing information while maintaining strong event generation quality.
This study addresses the representation, visualization, and quantification of regular structures in rhythmic data by proposing a “rhythmic motif” framework that models rhythms as fixed-length sequences of inter-onset intervals, decomposed into duration and pattern (the ratios between successive intervals). Building on this decomposition, the authors introduce pattern–duration plots and cluster transition networks to uncover rhythmic regularities. They reformulate the normalized Pairwise Variability Index (nPVI) as the mean deviation from isochrony and propose a more general measure of anisochrony alongside a novel concept termed “quantization.” This approach unifies existing visualization techniques and effectively reveals small-integer-ratio rhythmic structures in both synthetic and real-world datasets, establishing a more flexible and theoretically grounded framework for rhythm analysis.
This work addresses the longstanding disconnect between semantic understanding and high-fidelity numerical generation in time series modeling, where generative models often rely on superficial patterns while comprehension models struggle to produce precise values. To bridge this gap, we propose the first vision-centric unified framework that synergistically enhances both capabilities through three key innovations: a novel bidirectional lossless time-series-to-image mapping (Bi-TSI), an explicit comprehension-guided generation mechanism, and a multi-task joint training architecture. We further introduce the TSUMM-Suite benchmark, comprising six understanding and two generation tasks, to holistically evaluate model performance. Extensive experiments demonstrate that our approach significantly improves both semantic comprehension accuracy and numerical generation fidelity, establishing a new paradigm for multimodal time series modeling.
This work addresses the lack of a precise mathematical characterization of end-to-end input–output mappings in existing state space models (SSMs), such as S4D. By establishing a rigorous correspondence between S4D and analytically solvable networks of nonlinear oscillators, the authors embed S4D into a ring topology and derive, for the first time, an explicit operator expression for forward propagation. This derivation yields the first complete end-to-end analytical formulation for modern SSMs, revealing the underlying mechanism of information wave propagation through the ring structure and the interactions induced by nonlinear decoders. The resulting framework provides a clear physical interpretation and enhanced interpretability, and it generalizes across several mainstream SSM architectures.
本文提出了一种时序可分解图像表示方法(TDIR),通过正交子空间将历史照片分解为日期和内容组件,以保持时间结构并有效进行图像检索。