Score
Designs, implements, and evaluates methods for segmenting inputs into tokens or chunks and for grouping those chunks semantically to optimize downstream processing, storage, retrieval, and computational efficiency; includes selecting chunk sizes, chunking algorithms, and tokenization techniques and measuring token distributions, fidelity, and performance. Also includes modeling token-related costs and incentive structures (tokenomics) and analyzing tradeoffs among information retention, latency, and resource use when choosing tokenization or chunking strategies.
This study addresses the lack of systematic evaluation and inconsistent benchmarks in existing document chunking strategies for dense retrieval. The authors propose the first two-dimensional taxonomy that encompasses structural, semantic-aware, and large language model (LLM)-guided chunking approaches, along with embedding timing considerations. They establish a unified reproducible framework to comprehensively evaluate diverse strategies—including fixed-length, paragraph-level, LumberChunker, and Late Chunking—across both within-document and corpus-level retrieval tasks. Their findings reveal that structural chunking outperforms LLM-based methods in corpus retrieval, while LumberChunker achieves the best performance in within-document retrieval. Notably, contextualized chunking improves corpus retrieval effectiveness but degrades within-document performance, highlighting a task-dependent trade-off that informs optimal chunking selection.
Discrete tokenizers lack a systematic, cross-task survey. Method: We propose the first unified analytical framework covering generation, understanding, recommendation, and information retrieval; introduce a hierarchical decomposition paradigm for tokenizer submodules; establish a cross-task taxonomy; and conduct a horizontal comparison of representative approaches—including VQ-VAE, SoundStream, K-means tokenization, semantic hashing, and cross-modal alignment—through the lenses of information theory, representation learning, and structured modeling. Contribution/Results: We identify three core challenges: semantic alignment, cross-modal generalization, and the efficiency–accuracy trade-off. Furthermore, we deliver a reusable evaluation dimension matrix and an open challenge map, providing both theoretical foundations and practical guidelines for designing next-generation tokenizers that are robust, interpretable, and cross-modal.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
This study addresses the lack of comprehensive evaluations balancing effectiveness and system overhead in chunking strategies for long-document retrieval. We systematically benchmark eight strategies across multi-scale corpora, quantifying both retrieval performance and operational costs, including indexing throughput and latency. Our findings reveal that complex chunking methods do not necessarily outperform simpler baselines; high computational costs often yield inconsistent gains alongside significant operational disparities. Consequently, this work establishes chunking as a multi-objective design decision, emphasizing that optimal strategies are highly contingent upon specific models, datasets, and evaluation metrics. These insights provide empirical evidence to guide trade-off optimization in practical deployment scenarios, challenging the assumption that increased algorithmic complexity inherently translates to superior system-level utility in document retrieval pipelines.
This study investigates how the information granularity of data units (tokens) influences the scaling laws of language models, with a focus on the role of compression rate—defined as average bytes per token—in determining compute-optimal configurations. By training 988 latent tokenization models (BLTs) with flexible compression rates across model sizes from 50M to 7B parameters, and validating across multilingual and subword tokenization settings, the authors find that under compute-optimal conditions, model size should scale proportionally with data volume measured in bytes rather than tokens. Moreover, the optimal compression rate is substantially lower than that achieved by standard byte-pair encoding (BPE) and decreases further as available compute budget diminishes. These findings offer a byte-centric tokenization strategy for designing more efficient language models.
This work investigates how tokenizer training data scale (1 GB–900 GB) affects tokenization quality. Using systematic controlled experiments and causal attribution analysis on multi-scale real-world corpora with mainstream algorithms (e.g., BPE), we quantitatively identify a pronounced diminishing-return phenomenon: beyond 300 GB, improvements in key metrics—including BLEU and vocabulary coverage—fall below 0.5%. Further analysis pinpoints the pre-tokenization stage as the critical bottleneck causing performance saturation. Our findings empirically establish an effective upper bound on tokenizer training data size and provide actionable, data-efficient guidance for designing lightweight, high-performance tokenizers. This study fills a fundamental gap in understanding data efficiency for core NLP components.
This work addresses the limitations of traditional RAG systems, which rely on fixed chunking strategies ill-suited for diverse document structures and lack task-agnostic metrics to evaluate chunk quality. The authors propose the first adaptive chunking framework that dynamically selects the optimal chunking method based on document characteristics. They introduce five novel document-level intrinsic metrics—such as References Completeness and Intrachunk Cohesion—to guide chunking strategy selection without requiring downstream task feedback. The framework integrates an LLM-regex chunker and a recursive merging chunker, augmented with post-processing techniques. Evaluated across multiple domains without any model or prompt tuning, the approach improves QA accuracy from 62–64% to 72% and increases the number of correctly answered questions by over 30% (from 49 to 65).
This study investigates the design of effective text chunking strategies within the framework of the German Civil Code to enhance the performance of Retrieval-Augmented Generation (RAG) systems on legal question-answering tasks. The authors systematically evaluate a range of chunking approaches, including those based on legal structure (articles, paragraphs, sentences, and propositions), fixed-size windows, context-aware segmentation, semantic clustering, and hierarchical retrieval via RAPTOR. Experimental results demonstrate that strategies preserving the inherent legal structure achieve significantly higher recall than more complex semantic methods, while also offering superior efficiency in terms of query latency, index construction time, and storage overhead. These findings highlight a critical trade-off between semantic enrichment and computational cost in legal RAG applications.
This study addresses the lack of systematic evaluation of text chunking strategies in existing Retrieval-Augmented Generation (RAG) systems across diverse scenarios, which hinders objective assessment of their performance and applicability. For the first time, it conducts controlled, cross-task, and cross-data-type experiments within a unified framework to comparatively analyze mainstream chunking methods—including fixed-length, semantic, and emerging techniques. The findings reveal that most advanced chunking approaches exhibit limited generalizability, performing effectively only in specific contexts. Furthermore, the work quantifies the trade-offs between effectiveness and computational cost across different strategies and highlights critical yet often overlooked limitations of chunking as a preprocessing step. These insights provide empirical grounding for the design and optimization of RAG systems.
This work addresses the limitations of conventional text chunking methods in retrieval-augmented generation (RAG), which often compromise semantic integrity and suffer from a “measurement trap”: removing heading-chain prefixes drastically reduces annotator agreement, casting doubt on the validity of prevailing ablation-based evaluations. To overcome these issues, the authors propose a three-stage, large language model–free semantic chunking pipeline—comprising heading segmentation, semantic merging, and heading-chain prefixing—that leverages inherent document structure to construct contextually coherent chunks and enhance retrieval relevance. Evaluated on a production-scale Markdown knowledge base with 1,600 queries, the approach achieves a 23.8% relative improvement in MRR@5 (from 0.374 to 0.463) overall and an 11.7% gain (from 0.828 to 0.925) on the answerable subset, with inter-annotator agreement reaching Cohen’s κ = 0.45.
This work addresses the challenge of optimizing compression ratios in token-free hierarchical models with byte-level dynamic chunking by proposing an Adaptive Target Dynamic Chunking (ATDC) mechanism. ATDC introduces curriculum learning into dynamic chunking control for the first time, progressively increasing the target compression ratio from low to high during training to ensure stable optimization. The method models the evolution of chunking through Bytes Per Innermost Chunk (BPIC). Evaluated on the FineWeb-Edu 100B dataset, ATDC achieves bits-per-byte (BPB) performance comparable to both token-level and byte-level baselines while demonstrating markedly improved training stability and superior performance across multiple downstream tasks.
Retrieval-augmented generation (RAG) pipelines for code completion rely on chunking to segment source files into retrievable units, yet chunking strategies are typically adopted without empirical justification, and practitioner recommendations are notably inconsistent. We present a controlled empirical study isolating the effect of chunking on code completion quality by crossing four representative strategies (Function, Declaration, Sliding Window, and cAST) with four retrievers, five generators, and nine parameter configurations on two benchmarks (RepoEval and CrossCodeEval), totaling 864 experimental settings. Our results reveal that chunking strategy has a statistically significant effect on RAG-based code completion. Contrary to intuition, chunking based on functions underperforms all other strategies by 3.57--5.64 percentage points on RepoEval (Cliff's delta = -1.0), while the remaining chunking strategies perform comparably. Our further analysis demonstrates that this observation holds across all retriever--generator combinations. We also find that cross-file context length is the dominant parameter: doubling from 2,048 to 8,192 tokens yields up to 4.2 percentage points of improvement, whereas chunk size has a weaker, non-monotonic effect. On the cost--quality Pareto front, Sliding Window and cAST dominate both benchmarks; Function chunking is never Pareto-optimal.