semantic chunking

Design and implement methods that partition text or passage collections into semantically coherent, variable-length chunks by clustering embedding or similarity representations; determine chunk boundaries, assign passages to clusters, and produce cluster indexes or retrieval-ready chunk sets for downstream search or retrieval.

semanticchunking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing document chunking methods rely solely on semantic similarity while ignoring spatial layout, leading to suboptimal segmentation in complex documents (e.g., multi-column or image-text interleaved layouts) and poor controllability of chunk length for LLM input constraints. To address this, we propose a structure-aware adaptive chunking method: it jointly models textual bounding boxes, semantic embeddings (BERT/MPNet), and spatial relationships to construct a weighted heterogeneous graph, then applies spectral clustering for semantic-structural co-optimization; additionally, a dynamic length truncation strategy enforces strict token limits. This is the first document chunking framework that explicitly integrates spatial structure modeling with semantic coherence. Experiments demonstrate a 12.7% F1-score improvement on multi-layout benchmarks, 98.3% intra-chunk semantic consistency, and 100% compliance with prescribed token constraints.

Document ChunkingPositional InformationVariable Chunk Length

Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking

Jul 14, 2025
HT
Hai Toan Nguyen
🏛️ VNU University of Engineering and Technology

Traditional RAG systems rely on fixed-size text chunking, ignoring document structure and thereby causing semantic fragmentation and suboptimal retrieval relevance. To address this, we propose a hierarchical text segmentation and clustering-enhanced RAG framework. First, we perform structure-aware paragraph-level segmentation; then, we apply semantic clustering on paragraph embeddings to construct a dual-granularity vector index—comprising both paragraph-level and cluster-level representations. During retrieval, our method jointly leverages fine-grained paragraph matching and coarse-grained cluster-level semantic generalization, yielding structurally informed and semantically coherent retrieval units. Evaluated on three open-domain question answering benchmarks—NarrativeQA, QuALITY, and QASPER—our approach significantly outperforms standard chunking baselines, improving answer accuracy and contextual consistency. This work establishes a scalable, structure-aware paradigm for semantic retrieval in RAG systems.

Addressing limitations of traditional chunking methodsEnhancing retrieval precision with hierarchical segmentationImproving semantic coherence in RAG text chunks

Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis

May 27, 2025
SR
Sinchana Ramakanth Bhat
🏛️ Fraunhofer IAIS

Chunk size selection in long-document retrieval exhibits dataset- and embedding-model-dependent sensitivity, causing significant performance fluctuations. Method: We conduct a systematic empirical analysis across diverse datasets (short-form vs. long-form question answering) and state-of-the-art embedding models (e.g., Stella, Snowflake), quantifying model-specific chunk-size sensitivity for the first time. Contribution/Results: We identify an intrinsic trade-off: smaller chunks (64–128 tokens) yield superior performance on factoid QA, whereas larger chunks (512–1024 tokens) substantially improve retrieval accuracy for long-context queries. Based on these findings, we propose the “chunk–model–data” triadic co-adaptation principle—a principled, reproducible, and transferable framework for optimizing chunking strategies in long-document retrieval. This work provides actionable guidelines for aligning chunk granularity with model architecture and task characteristics, advancing robustness and generalizability in retrieval systems.

Analyzes chunk size impact on different embedding modelsEvaluates optimal chunk sizes for retrieval in diverse datasetsHighlights trade-offs between chunk size, models, and dataset types

To address lexical mismatch between queries and documents in information retrieval, existing query expansion methods suffer from context sensitivity and unstable performance, while document expansion approaches (e.g., Doc2Query) incur high preprocessing overhead, index bloat, and low generation reliability. This paper proposes a structured document enhancement paradigm: documents are segmented into semantically coherent chunks, and a single-encoder dual-decoder multitask learning framework—built upon T5—is trained to jointly generate chunk-level titles, candidate questions, and keywords in parallel. The generated metadata enriches retrieval inputs without modifying the underlying index structure. Evaluated on 305 query–document pairs, our method achieves 95.41% Top@10 accuracy, substantially outperforming baseline approaches. Results demonstrate that the proposed lightweight, reliable, and plug-and-play document enhancement solution effectively bridges lexical gaps while preserving efficiency and deployment flexibility.

Addresses vocabulary mismatch in information retrieval systemsGenerates structured semantic data from document chunksOvercomes document expansion limitations like preprocessing costs

To address the degradation in retrieval performance caused by contextual information loss in conventional text chunking for embedding, this paper proposes *late chunking*: chunking is performed after the top-layer Transformer output but before mean pooling, followed by token-level context-aware aggregation to generate chunk embeddings. This approach delays the chunking operation until after global contextual representations have been fully modeled, thereby inherently preserving long-range context without requiring model retraining—enabling plug-and-play adaptation to diverse long-context embedding models. Additionally, a lightweight fine-tuning strategy is introduced to enhance chunk-level discriminability. Extensive evaluation across multiple retrieval benchmarks demonstrates that late chunking significantly outperforms traditional early-chunking baselines, achieving consistent improvements in recall and relevance metrics. Crucially, the method maintains zero-shot deployability—requiring no training or architectural modification—while delivering state-of-the-art retrieval performance.

Chunk embeddings lack surrounding context, reducing effectiveness.Need method to embed long text before chunking.Retrieving smaller text portions loses contextual information.

Latest Papers

What's happening recently
View more

This study investigates optimal text chunking strategies for enhancing the response quality of Retrieval-Augmented Generation (RAG) systems when applied to structurally complex academic papers. We systematically compare semantic clustering, fixed-length, and recursive chunking approaches, evaluating output faithfulness and relevance using the RAGAs framework. To our knowledge, this is the first empirical comparison of multiple chunking strategies on long-form scholarly texts. Our findings indicate that semantic clustering does not significantly outperform simpler methods, and that question type—generic versus document-specific—substantially influences system performance. Furthermore, we identify limitations in the reliability of RAGAs’ faithfulness metric for such tasks, suggesting a need for more robust evaluation measures in academic RAG applications.

academic textschunking strategiesRAG evaluation

This study investigates the design of effective text chunking strategies within the framework of the German Civil Code to enhance the performance of Retrieval-Augmented Generation (RAG) systems on legal question-answering tasks. The authors systematically evaluate a range of chunking approaches, including those based on legal structure (articles, paragraphs, sentences, and propositions), fixed-size windows, context-aware segmentation, semantic clustering, and hierarchical retrieval via RAPTOR. Experimental results demonstrate that strategies preserving the inherent legal structure achieve significantly higher recall than more complex semantic methods, while also offering superior efficiency in terms of query latency, index construction time, and storage overhead. These findings highlight a critical trade-off between semantic enrichment and computational cost in legal RAG applications.

chunkingGerman statutory lawlegal information retrieval

This work addresses the limitations of traditional RAG systems, which rely on fixed chunking strategies ill-suited for diverse document structures and lack task-agnostic metrics to evaluate chunk quality. The authors propose the first adaptive chunking framework that dynamically selects the optimal chunking method based on document characteristics. They introduce five novel document-level intrinsic metrics—such as References Completeness and Intrachunk Cohesion—to guide chunking strategy selection without requiring downstream task feedback. The framework integrates an LLM-regex chunker and a recursive merging chunker, augmented with post-processing techniques. Evaluated across multiple domains without any model or prompt tuning, the approach improves QA accuracy from 62–64% to 72% and increases the number of correctly answered questions by over 30% (from 49 to 65).

chunkingdocument segmentationRAG

This work addresses the granularity mismatch between existing listwise explanation methods—which rely on isolated terms—and the semantic chunk representations leveraged by dense retrievers. To bridge this gap, the authors propose ChunkGroupSHAP, the first approach to incorporate semantic chunk clustering into Shapley value computation. By aligning attribution with the dense retriever’s representation granularity through cross-document semantic groupings, ChunkGroupSHAP enhances attribution consistency while preserving the listwise explanation framework. Experiments reveal that the optimal explanation unit varies with both the retriever and corpus: BM25 performs best with word-level units, dense models like E5 benefit from corpus-level groupings, and heterogeneous retrieval settings gain from query-local groupings. The method demonstrates consistent effectiveness across MS MARCO, FinanceBench, AILACaseDocs, and FinQA benchmarks.

dense rankersembedding-based rankingfeature-unit mismatch

Hot Scholars

TA

Tu Anh Dinh

Doctoral Researcher, Karlsruhe Institute of Technology
Data ScienceArtificial IntelligenceMachine Translation
DT

David T. Hoffmann

University of Freiburg
TransformersSelf-supervised LearningRepresentation LearningSynthetic Data
YF

Yixing Fan

ict
relevance rankingdeep learninginformation retrieval