Score
Designs, builds, and evaluates models and algorithms that represent, analyze, and generate temporal sequences at multiple temporal resolutions—combining hierarchies or U-shaped/multi-resolution architectures, dilated temporal convolutions, and transformer-based modules—to capture both short‑term fine-grained dynamics and long‑term cyclical or global patterns. Ensures preservation of frame‑level fidelity and temporal coherence (e.g., smooth, physically plausible trajectories, reduced jitter) by aggregating features across scales and enforcing local inter‑frame consistency.
This work addresses the lack of systematic evaluation of sequence models’ ability to capture diverse temporal dependencies—such as short- and long-range, decaying, and oscillatory patterns. We propose the first synthetic benchmark framework based on controllable, parameterized memory functions. By explicitly designing memory kernel functions, our framework generates synthetic tasks with continuous-time complexity, enabling fine-grained, interpretable, and theoretically grounded analysis of model memory characteristics. We evaluate mainstream architectures—including RNNs, Transformers, and State Space Models (SSMs)—under a unified benchmark across multiple dimensions. Our experiments not only validate existing theoretical predictions but also uncover, for the first time, implicit architectural preferences for specific memory patterns and their precise failure boundaries. The results provide reproducible, interpretable, and quantitative guidance for selecting appropriate sequence modeling paradigms.
Real-world multivariate time series exhibit strong non-stationarity and cross-scale dynamics; however, prevailing models rely on fixed-scale priors—such as chunk-based tokenization or static frequency-domain transformations—limiting modeling flexibility and hindering robustness to abrupt, high-magnitude events. To address this, we propose an adaptive hierarchical architecture that jointly captures instantaneous fluctuations and long-term trends via multi-scale convolutional encoding, integrates sequential modeling using BiLSTM or Transformer backbones, incorporates Squeeze-and-Excitation gating for channel-wise feature recalibration, and employs multi-head temporal attention for context-aware dynamic feature fusion. The resulting framework establishes a unified paradigm for time-series modeling. Evaluated across 32 benchmark datasets on forecasting, imputation, and classification tasks, our method achieves state-of-the-art performance on 24 datasets—outperforming leading approaches including EMTSF, TimesNet, and PatchTST by significant margins.
This work addresses the longstanding disconnect between semantic understanding and high-fidelity numerical generation in time series modeling, where generative models often rely on superficial patterns while comprehension models struggle to produce precise values. To bridge this gap, we propose the first vision-centric unified framework that synergistically enhances both capabilities through three key innovations: a novel bidirectional lossless time-series-to-image mapping (Bi-TSI), an explicit comprehension-guided generation mechanism, and a multi-task joint training architecture. We further introduce the TSUMM-Suite benchmark, comprising six understanding and two generation tasks, to holistically evaluate model performance. Extensive experiments demonstrate that our approach significantly improves both semantic comprehension accuracy and numerical generation fidelity, establishing a new paradigm for multimodal time series modeling.
Large language models (LLMs) inherently struggle to capture continuous temporal dynamics and explicit inter-variable dependencies in time series analysis. Method: This paper systematically reviews the emerging “time-series-to-image + vision model” paradigm, proposing the first dual-dimensional taxonomy: (i) time-series image encoding strategies (e.g., Gramian Angular Field, Markov Transition Field) and (ii) vision-model adaptation architectures (e.g., Vision Transformers, multimodal alignment, feature-decoupled reconstruction). It rigorously defines key pre-/post-processing challenges, surveys over 100 works, and establishes a unified evaluation framework. Contribution/Results: Empirical results demonstrate that vision-based approaches consistently outperform pure sequence models—achieving average accuracy gains of 5–12% across anomaly detection, forecasting, and classification tasks—thereby offering a promising new direction for time-series modeling.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
This work addresses the ambiguity regarding whether existing single-stage video object detectors genuinely leverage temporal context, as standard evaluation metrics often fail to reveal their actual reliance on temporal information. To this end, we propose TemporalLens, a diagnostic framework that quantifies a model’s temporal dependency through controlled perturbations—including temporal shuffling, structured occlusion, and redundancy injection. Furthermore, we design YOLO-3D based on YOLOv8, explicitly preserving the temporal dimension within the backbone to enhance genuine temporal reasoning. Experiments demonstrate that TemporalLens effectively distinguishes between stacked 2D models and true temporal architectures, while YOLO-3D achieves an average mAP@50 improvement of 3.7 percentage points with 32-frame inputs, underscoring the critical role of temporal depth in performance gains.
This work addresses the challenge of reconstructing high-temporal-resolution time series from low-resolution inputs, a task constrained by cost and feasibility in real-world data acquisition. To this end, the authors propose the SRT framework, which introduces decoupled rectified flows to temporal super-resolution for the first time. SRT decomposes input signals into trend and seasonal components and leverages implicit neural representations to align with the target resolution. A novel cross-resolution attention mechanism is designed to synthesize high-resolution details effectively. Building upon large-scale pretraining, the SRT-large variant demonstrates strong zero-shot super-resolution capabilities. Extensive experiments across nine public datasets show that SRT consistently outperforms existing methods, supports arbitrary scale factors, and validates both the efficacy of its individual components and the overall robustness of the framework.
This work addresses the inherent trade-off in existing state space models between preserving long-term memory and capturing high-resolution recent dynamics, a limitation rooted in the HiPPO operator’s conflict between time-scale invariance and sensitivity to local temporal variations. To resolve this, the study introduces fractional measure theory into state space modeling for the first time, proposing a recursive memory update architecture based on a tunable singularity-index projection operator. This design retains scale-invariant memory structures while enhancing responsiveness to recent perturbations. Furthermore, by diagonalizing the state space, the model enables joint multi-scale temporal feature learning. Evaluated on the Long Range Arena benchmark, the method achieves an average accuracy of 87.11%, including 61.85% on ListOps, outperforming state-of-the-art models such as S5.
This work addresses the tension between limited visual token budgets and the need for precise capture of critical events in long-form video multi-event understanding, as well as the lack of adaptive allocation and self-correction capabilities in existing two-stage approaches. To this end, we propose MoD-VLLM, a novel framework featuring a modular dynamic granularity mechanism and a reflexive closed-loop architecture. By iteratively co-optimizing temporal localization and semantic understanding through positive-negative segment localization and a dynamic granularity reflection module, MoD-VLLM enables fine-grained encoding of relevant segments and coarse-grained compression of irrelevant ones. A reinforcement learning strategy is further introduced to jointly optimize localization and representation. Extensive experiments on multiple long-video benchmarks and the newly constructed MEventBench demonstrate that MoD-VLLM significantly outperforms state-of-the-art methods, validating its effectiveness in complex multi-event reasoning.