Score
Designs and implements positional encoding schemes that inject absolute or relative event-time information into token embeddings by mapping timestamps or inter-event intervals to vector representations compatible with sequence models. Builds, configures, and evaluates learned or fixed time-aware encoding functions to represent irregular temporal gaps and to analyze their effect on the model’s temporal discrimination and downstream sequence performance.
Transformer-based time-series modeling suffers from positional encoding mismatches that impair temporal order representation. Method: We systematically survey, unify, and quantitatively benchmark time-series-specific positional encodings—including fixed, learnable, relative, and multi-scale hybrid variants—across standardized UCR/UEA datasets for time-series classification. We conduct cross-method and cross-task generalization analysis to characterize performance boundaries and applicability of each encoding paradigm. Contribution/Results: We propose design principles and improvement pathways tailored to time-series characteristics and release an open-source, standardized evaluation framework. Empirical results show that multi-scale hybrid encodings significantly enhance long-range dependency modeling, while learnable encodings exhibit superior robustness in low-data regimes. This work provides evidence-based guidance and practical recommendations for positional encoding selection and innovation in time-series Transformers.
This paper addresses the challenge of temporal representation in large language models (LLMs) for continuous-time event sequence modeling. We systematically evaluate five time tokenization strategies—byte encoding, adaptive residual scalar quantization, calendar-semantic formatting, uniform binning, and raw numeric string encoding—across diverse real-world event time distributions (e.g., log-normal, discrete spikes). To our knowledge, this is the first empirical study on time tokenization for LLMs. Results show that log-transformed tokenization achieves optimal performance on skewed distributions, whereas calendar-semantic formatting exhibits strongest robustness on multimodal and mixed distributions. Experiments, conducted within an LLM fine-tuning framework across multiple real-world event datasets, empirically validate the strategy–distribution alignment principle. We precisely delineate the performance boundaries: log-based tokenization excels on skewed distributions, while human-centric (calendar-semantic) tokenization dominates on mixed distributions. Our work establishes a reproducible, distribution-aware methodology for temporal representation in event sequence modeling.
Existing sequence models often neglect explicit temporal information, limiting their capacity for temporal reasoning and event timing modeling. This work proposes ChronoSSM, an autoregressive state space model that jointly models event content and timestamps within the SSM framework for the first time. By employing a shared backbone network, ChronoSSM simultaneously optimizes event prediction and time generation objectives, departing from conventional two-stage paradigms and enabling learned representations to encapsulate both semantic and temporal structures. Experiments across four datasets with varying levels of temporal annotation density demonstrate that ChronoSSM substantially improves the recoverability of inter-event timing information while maintaining strong event generation quality.
This work addresses text-driven Temporal Event Sequence Retrieval (TESR), aiming to accurately locate temporally ordered event sequences from natural language queries—enabling applications in e-commerce behavior analysis, social media monitoring, and criminal investigation. To advance this task, we introduce TESRBench, the first comprehensive benchmark encompassing diverse real-world, multi-source scenarios. We propose TPP-Embedding, the first unified framework integrating Large Language Models (LLMs) with Temporal Point Processes (TPPs), termed TPP-LLM. It incorporates sequence-level text-temporal joint contrastive learning, event-sequence pooling embeddings, and a hybrid text generation strategy combining synthetic data synthesis with human verification. Evaluated across the full TESRBench dataset, our method significantly outperforms existing baselines, achieving state-of-the-art performance in both retrieval accuracy and robustness.
Existing video generation models rely on single-paragraph textual prompts, making precise temporal control over multiple events challenging—often resulting in event omissions or chronological inconsistencies. To address this, we propose the first multi-event video generation framework enabling explicit specification of start and end time intervals for each event. Our method binds each event with fine-grained temporal annotations and introduces ReRoPE, a time-aware positional encoding, to achieve cross-modal temporal alignment between event descriptions and video tokens. Built upon a video diffusion Transformer architecture, the model is fine-tuned on temporally annotated video data. Experiments demonstrate significant improvements over state-of-the-art commercial and open-source models in event completeness, temporal accuracy, and transition naturalness. To our knowledge, this is the first approach enabling controllable spatiotemporal orchestration of multiple events within generated videos.
This work proposes a multimodal generative framework that reformulates time series classification as a text generation task, jointly modeling numerical sequences, textual context, and task instructions. Traditional approaches often struggle to incorporate contextual information and overlook semantic relationships among classes. To address these limitations, the framework employs time series discretization, an alignment projection layer, and generative self-supervised pretraining, complemented by an implicit feature augmentation mechanism that integrates statistical features with vision-language image descriptions. This design effectively compensates for the inductive bias deficiencies of language models in temporal modeling. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method significantly outperforms existing approaches, highlighting its superior performance and strong generalization capability.
Speculative decoding (SD) accelerates large language model (LLM) inference but is constrained by the requirement that the draft and target models share an identical vocabulary—limiting draft model selection and often necessitating costly retraining. This work proposes TokenTiming, the first SD framework to integrate dynamic time warping (DTW) for cross-vocabulary speculative decoding. TokenTiming dynamically aligns token sequences and probability distributions via sequence recoding and DTW-based soft alignment, enabling seamless cooperation between arbitrary off-the-shelf models without architectural modification or retraining. Crucially, it eliminates vocabulary compatibility constraints while preserving decoding correctness and efficiency. Experiments across diverse NLP tasks demonstrate an average 1.57× inference speedup over standard autoregressive decoding, with consistent latency reduction and throughput improvement. TokenTiming significantly enhances the practicality, flexibility, and generalizability of speculative decoding, establishing a foundation for vocabulary-agnostic acceleration of LLM inference.
Existing video generation models struggle to model overlapping events because their temporal representations are limited to discrete time steps, lacking the ability to capture time intervals and concurrent relationships. This work proposes Temporal Interval Encoding (TIE), which for the first time treats time intervals as first-class primitives within a RoPE-compatible bilinear attention mechanism. By leveraging interval integration and duration invariance principles, TIE derives a closed-form sinc-based solution that enables interval-aware generation without altering the standard attention interface. Evaluated on the OmniEvents dataset, TIE improves the human-verified temporal constraint satisfaction rate from 77.34% to 96.03% and reduces temporal boundary error from 0.261 seconds to 0.073 seconds, all while preserving the original visual quality of the DiT architecture.
Transformer models are inherently insensitive to word order and rely on positional encodings to inject sequential information, yet existing designs often lack a rigorous theoretical foundation. This work proposes a geometric framework for positional encoding, establishing its necessity and separability, and derives a minimally parameterized representation. Building upon the Hellinger distance and classical multidimensional scaling (MDS), the authors construct an information-theoretically optimal encoding scheme. By leveraging matrix rank analysis and neural tangent kernel (NTK) theory, they unify the evaluation of encoding quality into a single stress metric. Empirical validation on SST-2 and IMDB demonstrates that ALiBi encodings exhibit significantly lower stress compared to sinusoidal and RoPE encodings, corroborating their near rank-1 optimal structure.
Transformer-based time series models face a trade-off between computational efficiency and information fidelity under fixed chunking strategies. This work proposes a content-aware dynamic chunking mechanism that adaptively adjusts chunk boundaries based on local signal complexity, enabling fine-grained representation in information-dense regions and coarse-grained aggregation in smooth segments. The approach integrates a lightweight state-space encoder with a dynamic chunking algorithm to deliver compressed yet informative temporal inputs to the Transformer. During large-scale pretraining, the model achieves up to 20× faster convergence and 8× improved data efficiency compared to baseline methods, while setting new state-of-the-art results on long-horizon forecasting benchmarks.