Score
Designs and implements methods that partition ordered data (sequences or time series) into contiguous patches and convert each patch into a discrete or continuous token representation suitable for transformer-style sequence models. This work involves choosing patch size, stride and overlap, segmentation algorithms and patch encoders, and measuring tokenization quality via metrics and experiments to improve learning efficiency, convergence, and reconstruction error relative to simple tokenization schemes.
Discrete tokenizers lack a systematic, cross-task survey. Method: We propose the first unified analytical framework covering generation, understanding, recommendation, and information retrieval; introduce a hierarchical decomposition paradigm for tokenizer submodules; establish a cross-task taxonomy; and conduct a horizontal comparison of representative approaches—including VQ-VAE, SoundStream, K-means tokenization, semantic hashing, and cross-modal alignment—through the lenses of information theory, representation learning, and structured modeling. Contribution/Results: We identify three core challenges: semantic alignment, cross-modal generalization, and the efficiency–accuracy trade-off. Furthermore, we deliver a reusable evaluation dimension matrix and an open challenge map, providing both theoretical foundations and practical guidelines for designing next-generation tokenizers that are robust, interpretable, and cross-modal.
This work addresses the limitations of Transformer-based models in long-sequence multivariate time series forecasting, particularly their inadequate input representation quality and structural modeling capacity. To overcome these challenges, the authors propose a two-stage framework: first, a convolutional neural network (CNN) extracts local dynamic features from fixed-length temporal segments and generates compact patch-level token embeddings; subsequently, a Transformer encoder with attention mechanisms models the global dependencies among these segments. By decoupling local feature extraction from global dependency modeling, the approach enhances both scalability and representational power. Experimental results on synthetic multivariate time series datasets demonstrate that the proposed method significantly outperforms CNN baselines under long input sequences and achieves performance comparable to state-of-the-art patch-based Transformer models.
To address the high computational overhead of Transformers and state-space models (SSMs) when modeling long time series, this paper introduces token merging—previously unexplored in time-series analysis—for the first time. We propose a domain-adapted local merging paradigm that enforces subsequence neighborhood constraints, employs linear weighted aggregation, and integrates lightweight attention/SSM modules to jointly preserve local dependency modeling and computational efficiency. Our method achieves up to 5400% inference speedup on state-of-the-art time-series foundation models (e.g., Chronos), with negligible accuracy degradation. It demonstrates robust performance across diverse architectures and benchmark datasets, significantly improving throughput for long sequences. This work establishes a novel pathway toward efficient deployment of large-scale time-series models.
Why do large language models (LLMs) require tokenization, and why does character-level modeling lead to performance degradation in Transformers? Method: The authors construct a *k*-order Markov data source and rigorously analyze the cross-entropy of Transformers under character-level versus token-level modeling, grounding the analysis in information-theoretic modeling capacity. They establish a provable relationship between tokenization strategies and the accuracy of sequence probability estimation. Contribution/Results: Theoretically, without tokenization, Transformers collapse to modeling only unigram character distributions, failing to capture higher-order dependencies; with appropriate tokenization, learning single-step token predictions suffices to near-optimally model the source distribution. Empirically, tokenization significantly reduces cross-entropy on high-order Markov sources. This work provides the first rigorous information-theoretic and probabilistic justification that tokenization is a necessary condition for overcoming the fundamental limitations of character-level Transformer modeling.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
Existing image tokenizers employ raster-scan spatial token ordering, which is inherently incompatible with autoregressive modeling. This work proposes a spectral-domain image tokenizer based on the discrete wavelet transform (DWT), mapping images into a coarse-to-fine, multi-scale spectral token sequence—establishing the first spectral-domain, coarse-grained-first tokenization paradigm. The method naturally enables zero-shot cross-resolution reconstruction, partial decoding for rapid coarse preview generation, text-guided super-resolution and editing, and adapts to arbitrary resolutions without retraining. Experiments demonstrate substantial improvements in token reconstruction fidelity and achieve state-of-the-art performance on multi-scale image generation, text-guided super-resolution, and text-guided image editing tasks.
This work investigates how different sequence representations—such as bytes, characters, and subwords—affect the information acquisition capacity of Transformer models under a fixed context window, a question that remains poorly understood. From an information-theoretic perspective, the paper introduces the notion of “fragmentation” and formally demonstrates that it inherently increases the log-loss of the optimal finite-context model. It establishes theoretical guarantees linking tokenization compression rates to the reliability of source context coverage, thereby constructing the first information-theoretic framework for representation selection in finite-context settings. Through Markov source modeling and comparative analysis of various tokenization strategies—including BPE, WordPiece, and byte-level methods—the study reveals the theoretical underpinnings of performance differences observed in models like ByT5 and CANINE, and proposes practical metrics to evaluate the effective context coverage of real-world tokenizers.
This study addresses the computational bottlenecks of byte-level language models caused by excessively long sequences and the absence of explicit textual abstraction. To overcome these limitations, this work proposes a tokenizer-free architecture that leverages a Token Hyper-position training strategy alongside hash embeddings, enabling standard Transformers to model efficiently at the byte level. The findings demonstrate that additional computation can effectively substitute fixed tokenizers, allowing models to spontaneously construct local context representation mechanisms. Furthermore, the proposed approach surpasses subword-based models in performance at large scales. Notably, by exploiting the non-uniform uncertainty inherent in generated outputs, the method facilitates speculative decoding, achieving a 3.4-fold improvement in acceptance rates.
This study addresses the computational inefficiency in time series language models arising from the unified treatment of time series and prompt tokens, which exhibit fundamentally different information structures. The work reveals an asymmetry in token importance: time series tokens contribute unevenly across the frequency spectrum, while the influence of prompt tokens diminishes with model depth. Leveraging this insight, the authors propose an adaptive token compression framework that dynamically compresses time series tokens via spectral analysis and progressively prunes prompt tokens in deeper layers, enabling hierarchical, frequency-aware, non-uniform budget allocation. Evaluated across forecasting, classification, imputation, and anomaly detection tasks, the method achieves up to 7.68× speedup and improves performance in 78% of experimental settings.
本文提出了一种通过合并模块压缩输入序列的方法,以减少Transformer模型的计算成本,同时保持准确性。
This study addresses the theoretical gap and generalization challenges in uniformly approximating causal mappings over arbitrarily long sequences using Transformers. To this end, this work characterizes cross-resolution causal families via α-Hölder continuity, establishing a universal approximation theory free of length-dependent parameters. By integrating masked attention mechanisms, it further derives generalization bounds within the infinite-length mean-field limit. The key contributions include achieving quantitative approximation without maximum-length factors, establishing a generalization bound of O((log log N/log N)^{β/(d+2)}), and validating the theoretical findings through experiments on physical time series data.