patch-based tokenization

Designs and implements methods that partition ordered data (sequences or time series) into contiguous patches and convert each patch into a discrete or continuous token representation suitable for transformer-style sequence models. This work involves choosing patch size, stride and overlap, segmentation algorithms and patch encoders, and measuring tokenization quality via metrics and experiments to improve learning efficiency, convergence, and reconstruction error relative to simple tokenization schemes.

patch-basedtokenization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of Transformer-based models in long-sequence multivariate time series forecasting, particularly their inadequate input representation quality and structural modeling capacity. To overcome these challenges, the authors propose a two-stage framework: first, a convolutional neural network (CNN) extracts local dynamic features from fixed-length temporal segments and generates compact patch-level token embeddings; subsequently, a Transformer encoder with attention mechanisms models the global dependencies among these segments. By decoupling local feature extraction from global dependency modeling, the approach enhances both scalability and representational power. Experimental results on synthetic multivariate time series datasets demonstrate that the proposed method significantly outperforms CNN baselines under long input sequences and achieves performance comparable to state-of-the-art patch-based Transformer models.

input representationmultivariate time-seriessequence length

Efficient Time Series Processing for Transformers and State-Space Models through Token Merging

May 28, 2024
LG
Leon Götz
🏛️ Volkswagen AG | Technical University of Munich

To address the high computational overhead of Transformers and state-space models (SSMs) when modeling long time series, this paper introduces token merging—previously unexplored in time-series analysis—for the first time. We propose a domain-adapted local merging paradigm that enforces subsequence neighborhood constraints, employs linear weighted aggregation, and integrates lightweight attention/SSM modules to jointly preserve local dependency modeling and computational efficiency. Our method achieves up to 5400% inference speedup on state-of-the-art time-series foundation models (e.g., Chronos), with negligible accuracy degradation. It demonstrates robust performance across diverse architectures and benchmark datasets, significantly improving throughput for long sequences. This work establishes a novel pathway toward efficient deployment of large-scale time-series models.

Developing local merging for scalable and causal token reductionEfficient processing of long token sequences in time series analysisPredicting merging benefits via spectral properties without task evaluation

Toward a Theory of Tokenization in LLMs

Apr 12, 2024
NR
Nived Rajaraman
🏛️ University of California, Berkeley

Why do large language models (LLMs) require tokenization, and why does character-level modeling lead to performance degradation in Transformers? Method: The authors construct a *k*-order Markov data source and rigorously analyze the cross-entropy of Transformers under character-level versus token-level modeling, grounding the analysis in information-theoretic modeling capacity. They establish a provable relationship between tokenization strategies and the accuracy of sequence probability estimation. Contribution/Results: Theoretically, without tokenization, Transformers collapse to modeling only unigram character distributions, failing to capture higher-order dependencies; with appropriate tokenization, learning single-step token predictions suffices to near-optimally model the source distribution. Empirically, tokenization significantly reduces cross-entropy on high-order Markov sources. This work provides the first rigorous information-theoretic and probabilistic justification that tokenization is a necessary condition for overcoming the fundamental limitations of character-level Transformer modeling.

Compare transformers' learning with and without tokenization on simple dataJustify tokenization use by analyzing cross-entropy loss in transformersStudy tokenization's role in transformer performance on Markov processes

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Spectral Image Tokenizer

Dec 12, 2024
CE
Carlos Esteves
🏛️ Google Research

Existing image tokenizers employ raster-scan spatial token ordering, which is inherently incompatible with autoregressive modeling. This work proposes a spectral-domain image tokenizer based on the discrete wavelet transform (DWT), mapping images into a coarse-to-fine, multi-scale spectral token sequence—establishing the first spectral-domain, coarse-grained-first tokenization paradigm. The method naturally enables zero-shot cross-resolution reconstruction, partial decoding for rapid coarse preview generation, text-guided super-resolution and editing, and adapts to arbitrary resolutions without retraining. Experiments demonstrate substantial improvements in token reconstruction fidelity and achieve state-of-the-art performance on multi-scale image generation, text-guided super-resolution, and text-guided image editing tasks.

Enables multi-resolution image handling without retrainingEnhances autoregressive modeling for image generationImproves image tokenization via spectral methods

Latest Papers

What's happening recently
View more

This work investigates how different sequence representations—such as bytes, characters, and subwords—affect the information acquisition capacity of Transformer models under a fixed context window, a question that remains poorly understood. From an information-theoretic perspective, the paper introduces the notion of “fragmentation” and formally demonstrates that it inherently increases the log-loss of the optimal finite-context model. It establishes theoretical guarantees linking tokenization compression rates to the reliability of source context coverage, thereby constructing the first information-theoretic framework for representation selection in finite-context settings. Through Markov source modeling and comparative analysis of various tokenization strategies—including BPE, WordPiece, and byte-level methods—the study reveals the theoretical underpinnings of performance differences observed in models like ByT5 and CANINE, and proposes practical metrics to evaluate the effective context coverage of real-world tokenizers.

finite-context predictionfragmentationrepresentation

This study addresses the computational bottlenecks of byte-level language models caused by excessively long sequences and the absence of explicit textual abstraction. To overcome these limitations, this work proposes a tokenizer-free architecture that leverages a Token Hyper-position training strategy alongside hash embeddings, enabling standard Transformers to model efficiently at the byte level. The findings demonstrate that additional computation can effectively substitute fixed tokenizers, allowing models to spontaneously construct local context representation mechanisms. Furthermore, the proposed approach surpasses subword-based models in performance at large scales. Notably, by exploiting the non-uniform uncertainty inherent in generated outputs, the method facilitates speculative decoding, achieving a 3.4-fold improvement in acceptance rates.

Byte Language ModelsEmergent AbstractionsTokenizer-free

This study addresses the computational inefficiency in time series language models arising from the unified treatment of time series and prompt tokens, which exhibit fundamentally different information structures. The work reveals an asymmetry in token importance: time series tokens contribute unevenly across the frequency spectrum, while the influence of prompt tokens diminishes with model depth. Leveraging this insight, the authors propose an adaptive token compression framework that dynamically compresses time series tokens via spectral analysis and progressively prunes prompt tokens in deeper layers, enabling hierarchical, frequency-aware, non-uniform budget allocation. Evaluated across forecasting, classification, imputation, and anomaly detection tasks, the method achieves up to 7.68× speedup and improves performance in 78% of experimental settings.

asymmetric tokenstime series language modelstoken compression

This study addresses the theoretical gap and generalization challenges in uniformly approximating causal mappings over arbitrarily long sequences using Transformers. To this end, this work characterizes cross-resolution causal families via α-Hölder continuity, establishing a universal approximation theory free of length-dependent parameters. By integrating masked attention mechanisms, it further derives generalization bounds within the infinite-length mean-field limit. The key contributions include achieving quantitative approximation without maximum-length factors, establishing a generalization bound of O((log log N/log N)^{β/(d+2)}), and validating the theoretical findings through experiments on physical time series data.

Causal TransformersContext Length GeneralizationGeneralization Bound

Hot Scholars

PL

Pengfei Liu

Associate professor at Shanghai Jiao Tong University
LLM
TP

Themis Palpanas

Distinguished Professor, University Paris Cite, French University Institute (IUF)
data managementdata sciencedata/time seriesanomaly detection
AB

Angela Bonifati

Distinguished Professor of Computer Science, Lyon 1 University, Senior Member of the French IUF
Graph DatabasesData IntegrationArtificial IntelligenceBig Data Analytics
RA

Ryan A. Rossi

Adobe Research
Machine LearningPersonalizationGraph Representation LearningGraph ML