Score
Design and implement methods that partition text or passage collections into semantically coherent, variable-length chunks by clustering embedding or similarity representations; determine chunk boundaries, assign passages to clusters, and produce cluster indexes or retrieval-ready chunk sets for downstream search or retrieval.
This study addresses the lack of systematic evaluation and inconsistent benchmarks in existing document chunking strategies for dense retrieval. The authors propose the first two-dimensional taxonomy that encompasses structural, semantic-aware, and large language model (LLM)-guided chunking approaches, along with embedding timing considerations. They establish a unified reproducible framework to comprehensively evaluate diverse strategies—including fixed-length, paragraph-level, LumberChunker, and Late Chunking—across both within-document and corpus-level retrieval tasks. Their findings reveal that structural chunking outperforms LLM-based methods in corpus retrieval, while LumberChunker achieves the best performance in within-document retrieval. Notably, contextualized chunking improves corpus retrieval effectiveness but degrades within-document performance, highlighting a task-dependent trade-off that informs optimal chunking selection.
Existing document chunking methods rely solely on semantic similarity while ignoring spatial layout, leading to suboptimal segmentation in complex documents (e.g., multi-column or image-text interleaved layouts) and poor controllability of chunk length for LLM input constraints. To address this, we propose a structure-aware adaptive chunking method: it jointly models textual bounding boxes, semantic embeddings (BERT/MPNet), and spatial relationships to construct a weighted heterogeneous graph, then applies spectral clustering for semantic-structural co-optimization; additionally, a dynamic length truncation strategy enforces strict token limits. This is the first document chunking framework that explicitly integrates spatial structure modeling with semantic coherence. Experiments demonstrate a 12.7% F1-score improvement on multi-layout benchmarks, 98.3% intra-chunk semantic consistency, and 100% compliance with prescribed token constraints.
Traditional RAG systems rely on fixed-size text chunking, ignoring document structure and thereby causing semantic fragmentation and suboptimal retrieval relevance. To address this, we propose a hierarchical text segmentation and clustering-enhanced RAG framework. First, we perform structure-aware paragraph-level segmentation; then, we apply semantic clustering on paragraph embeddings to construct a dual-granularity vector index—comprising both paragraph-level and cluster-level representations. During retrieval, our method jointly leverages fine-grained paragraph matching and coarse-grained cluster-level semantic generalization, yielding structurally informed and semantically coherent retrieval units. Evaluated on three open-domain question answering benchmarks—NarrativeQA, QuALITY, and QASPER—our approach significantly outperforms standard chunking baselines, improving answer accuracy and contextual consistency. This work establishes a scalable, structure-aware paradigm for semantic retrieval in RAG systems.
Chunk size selection in long-document retrieval exhibits dataset- and embedding-model-dependent sensitivity, causing significant performance fluctuations. Method: We conduct a systematic empirical analysis across diverse datasets (short-form vs. long-form question answering) and state-of-the-art embedding models (e.g., Stella, Snowflake), quantifying model-specific chunk-size sensitivity for the first time. Contribution/Results: We identify an intrinsic trade-off: smaller chunks (64–128 tokens) yield superior performance on factoid QA, whereas larger chunks (512–1024 tokens) substantially improve retrieval accuracy for long-context queries. Based on these findings, we propose the “chunk–model–data” triadic co-adaptation principle—a principled, reproducible, and transferable framework for optimizing chunking strategies in long-document retrieval. This work provides actionable guidelines for aligning chunk granularity with model architecture and task characteristics, advancing robustness and generalizability in retrieval systems.
To address lexical mismatch between queries and documents in information retrieval, existing query expansion methods suffer from context sensitivity and unstable performance, while document expansion approaches (e.g., Doc2Query) incur high preprocessing overhead, index bloat, and low generation reliability. This paper proposes a structured document enhancement paradigm: documents are segmented into semantically coherent chunks, and a single-encoder dual-decoder multitask learning framework—built upon T5—is trained to jointly generate chunk-level titles, candidate questions, and keywords in parallel. The generated metadata enriches retrieval inputs without modifying the underlying index structure. Evaluated on 305 query–document pairs, our method achieves 95.41% Top@10 accuracy, substantially outperforming baseline approaches. Results demonstrate that the proposed lightweight, reliable, and plug-and-play document enhancement solution effectively bridges lexical gaps while preserving efficiency and deployment flexibility.
To address the degradation in retrieval performance caused by contextual information loss in conventional text chunking for embedding, this paper proposes *late chunking*: chunking is performed after the top-layer Transformer output but before mean pooling, followed by token-level context-aware aggregation to generate chunk embeddings. This approach delays the chunking operation until after global contextual representations have been fully modeled, thereby inherently preserving long-range context without requiring model retraining—enabling plug-and-play adaptation to diverse long-context embedding models. Additionally, a lightweight fine-tuning strategy is introduced to enhance chunk-level discriminability. Extensive evaluation across multiple retrieval benchmarks demonstrates that late chunking significantly outperforms traditional early-chunking baselines, achieving consistent improvements in recall and relevance metrics. Crucially, the method maintains zero-shot deployability—requiring no training or architectural modification—while delivering state-of-the-art retrieval performance.
This study investigates optimal text chunking strategies for enhancing the response quality of Retrieval-Augmented Generation (RAG) systems when applied to structurally complex academic papers. We systematically compare semantic clustering, fixed-length, and recursive chunking approaches, evaluating output faithfulness and relevance using the RAGAs framework. To our knowledge, this is the first empirical comparison of multiple chunking strategies on long-form scholarly texts. Our findings indicate that semantic clustering does not significantly outperform simpler methods, and that question type—generic versus document-specific—substantially influences system performance. Furthermore, we identify limitations in the reliability of RAGAs’ faithfulness metric for such tasks, suggesting a need for more robust evaluation measures in academic RAG applications.
This study investigates the design of effective text chunking strategies within the framework of the German Civil Code to enhance the performance of Retrieval-Augmented Generation (RAG) systems on legal question-answering tasks. The authors systematically evaluate a range of chunking approaches, including those based on legal structure (articles, paragraphs, sentences, and propositions), fixed-size windows, context-aware segmentation, semantic clustering, and hierarchical retrieval via RAPTOR. Experimental results demonstrate that strategies preserving the inherent legal structure achieve significantly higher recall than more complex semantic methods, while also offering superior efficiency in terms of query latency, index construction time, and storage overhead. These findings highlight a critical trade-off between semantic enrichment and computational cost in legal RAG applications.
This work addresses the limitations of traditional RAG systems, which rely on fixed chunking strategies ill-suited for diverse document structures and lack task-agnostic metrics to evaluate chunk quality. The authors propose the first adaptive chunking framework that dynamically selects the optimal chunking method based on document characteristics. They introduce five novel document-level intrinsic metrics—such as References Completeness and Intrachunk Cohesion—to guide chunking strategy selection without requiring downstream task feedback. The framework integrates an LLM-regex chunker and a recursive merging chunker, augmented with post-processing techniques. Evaluated across multiple domains without any model or prompt tuning, the approach improves QA accuracy from 62–64% to 72% and increases the number of correctly answered questions by over 30% (from 49 to 65).
This work addresses the granularity mismatch between existing listwise explanation methods—which rely on isolated terms—and the semantic chunk representations leveraged by dense retrievers. To bridge this gap, the authors propose ChunkGroupSHAP, the first approach to incorporate semantic chunk clustering into Shapley value computation. By aligning attribution with the dense retriever’s representation granularity through cross-document semantic groupings, ChunkGroupSHAP enhances attribution consistency while preserving the listwise explanation framework. Experiments reveal that the optimal explanation unit varies with both the retriever and corpus: BM25 performs best with word-level units, dense models like E5 benefit from corpus-level groupings, and heterogeneous retrieval settings gain from query-local groupings. The method demonstrates consistent effectiveness across MS MARCO, FinanceBench, AILACaseDocs, and FinQA benchmarks.