Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis

๐Ÿ“… 2025-05-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Chunk size selection in long-document retrieval exhibits dataset- and embedding-model-dependent sensitivity, causing significant performance fluctuations. Method: We conduct a systematic empirical analysis across diverse datasets (short-form vs. long-form question answering) and state-of-the-art embedding models (e.g., Stella, Snowflake), quantifying model-specific chunk-size sensitivity for the first time. Contribution/Results: We identify an intrinsic trade-off: smaller chunks (64โ€“128 tokens) yield superior performance on factoid QA, whereas larger chunks (512โ€“1024 tokens) substantially improve retrieval accuracy for long-context queries. Based on these findings, we propose the โ€œchunkโ€“modelโ€“dataโ€ triadic co-adaptation principleโ€”a principled, reproducible, and transferable framework for optimizing chunking strategies in long-document retrieval. This work provides actionable guidelines for aligning chunk granularity with model architecture and task characteristics, advancing robustness and generalizability in retrieval systems.

Technology Category

Data Mining & Knowledge Management: Conversational Systems for Recommendation & RetrievalNatural Language Processing: Syntax โ€” Tagging, Chunking & ParsingSearch and Optimization: Evaluation and Analysis

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphs
๐Ÿ“ Abstract
Chunking is a crucial preprocessing step in retrieval-augmented generation (RAG) systems, significantly impacting retrieval effectiveness across diverse datasets. In this study, we systematically evaluate fixed-size chunking strategies and their influence on retrieval performance using multiple embedding models. Our experiments, conducted on both short-form and long-form datasets, reveal that chunk size plays a critical role in retrieval effectiveness -- smaller chunks (64-128 tokens) are optimal for datasets with concise, fact-based answers, whereas larger chunks (512-1024 tokens) improve retrieval in datasets requiring broader contextual understanding. We also analyze the impact of chunking on different embedding models, finding that they exhibit distinct chunking sensitivities. While models like Stella benefit from larger chunks, leveraging global context for long-range retrieval, Snowflake performs better with smaller chunks, excelling at fine-grained, entity-based matching. Our results underscore the trade-offs between chunk size, embedding models, and dataset characteristics, emphasizing the need for improved chunk quality measures, and more comprehensive datasets to advance chunk-based retrieval in long-document Information Retrieval (IR).
Problem

Research questions and friction points this paper is trying to address.

Evaluates optimal chunk sizes for retrieval in diverse datasets
Analyzes chunk size impact on different embedding models
Highlights trade-offs between chunk size, models, and dataset types
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates fixed-size chunking strategies for RAG
Optimal chunk size varies by dataset type
Embedding models show distinct chunking sensitivities
๐Ÿ”Ž Similar Papers
S
Sinchana Ramakanth Bhat
Fraunhofer IAIS, Germany
M
Max Rudat
Fraunhofer IAIS, Germany
J
Jannis Spiekermann
Fraunhofer IAIS, Germany
N
Nicolas Flores-Herr
Fraunhofer IAIS, Germany