Score
Design and implement pipelines that split long documents into independent chunks, encode each chunk into vectors (optionally with a frozen encoder), and combine those chunk vectors into a single document-level representation, score, or inference output. This includes methods for chunk-based evidence aggregation and document inference that preserve a one-query–one-document interface while reducing evidence dilution across long texts.
Long-document retrieval faces core challenges including excessive length, scattered evidence, and complex structural organization. This paper systematically surveys the evolution of long-document retrieval techniques across the pre-trained model and large language model (LLM) eras, unifying three developmental stages—classical passage retrieval, hierarchical encoding with efficient attention mechanisms, and LLM-driven re-ranking—for the first time. We propose a unified technical taxonomy encompassing key paradigms such as passage aggregation, structure-aware encoding, and retrieval-augmented generation, alongside domain-specific evaluation resources. Furthermore, we identify critical open problems in the foundation model era: the efficiency–effectiveness trade-off, cross-modal semantic alignment, and result interpretability. This work establishes an authoritative, structured survey framework and a comprehensive roadmap for long-document information retrieval research.
This work addresses the performance degradation in dense retrieval of long documents, where full-document encoding dilutes critical information. The authors propose DICE, a method that preserves the standard single-query–single-document interface while independently encoding document chunks and aggregating them into a single vector to retain the strongest evidence signals. They introduce the Evidence Dilution Index (EDI) to quantify information loss in document representations and design a training-free, document-side chunk aggregation strategy. Evaluated on LongEmbed using frozen dense retrieval models, DICE substantially improves long-document retrieval: for passages longer than 4k tokens, Passkey and Needle task scores rise from 30.0/23.3 to 90.0/74.0, with 92.8% of samples exhibiting significantly reduced EDI.
To address severe information loss in hundred-document-scale multi-document event summarization, this work systematically compares compression-based (multi-stage pipeline) and full-text-based (direct long-context modeling) approaches. Leveraging long-context Transformers—including Llama-3.1, Command-R, and Jamba-1.5-Mini—augmented with retrieval enhancement, hierarchical compression, and incremental summarization, we find that full-text modeling combined with retrieval achieves the best overall performance; in contrast, compression methods better preserve local information at intermediate stages but suffer from global context loss and consequent information decay. Building on these insights, we propose a hybrid paradigm that synergistically integrates compression and full-text modeling at critical stages. Empirical evaluation demonstrates that this architecture significantly improves summary coherence, coverage, and factual consistency. The approach offers a scalable, principled solution for large-scale event summarization, advancing the state of the art in handling ultra-long document collections.
This study addresses the lack of systematic evaluation and inconsistent benchmarks in existing document chunking strategies for dense retrieval. The authors propose the first two-dimensional taxonomy that encompasses structural, semantic-aware, and large language model (LLM)-guided chunking approaches, along with embedding timing considerations. They establish a unified reproducible framework to comprehensively evaluate diverse strategies—including fixed-length, paragraph-level, LumberChunker, and Late Chunking—across both within-document and corpus-level retrieval tasks. Their findings reveal that structural chunking outperforms LLM-based methods in corpus retrieval, while LumberChunker achieves the best performance in within-document retrieval. Notably, contextualized chunking improves corpus retrieval effectiveness but degrades within-document performance, highlighting a task-dependent trade-off that informs optimal chunking selection.
To address lexical mismatch between queries and documents in information retrieval, existing query expansion methods suffer from context sensitivity and unstable performance, while document expansion approaches (e.g., Doc2Query) incur high preprocessing overhead, index bloat, and low generation reliability. This paper proposes a structured document enhancement paradigm: documents are segmented into semantically coherent chunks, and a single-encoder dual-decoder multitask learning framework—built upon T5—is trained to jointly generate chunk-level titles, candidate questions, and keywords in parallel. The generated metadata enriches retrieval inputs without modifying the underlying index structure. Evaluated on 305 query–document pairs, our method achieves 95.41% Top@10 accuracy, substantially outperforming baseline approaches. Results demonstrate that the proposed lightweight, reliable, and plug-and-play document enhancement solution effectively bridges lexical gaps while preserving efficiency and deployment flexibility.
To address the degradation in retrieval performance caused by contextual information loss in conventional text chunking for embedding, this paper proposes *late chunking*: chunking is performed after the top-layer Transformer output but before mean pooling, followed by token-level context-aware aggregation to generate chunk embeddings. This approach delays the chunking operation until after global contextual representations have been fully modeled, thereby inherently preserving long-range context without requiring model retraining—enabling plug-and-play adaptation to diverse long-context embedding models. Additionally, a lightweight fine-tuning strategy is introduced to enhance chunk-level discriminability. Extensive evaluation across multiple retrieval benchmarks demonstrates that late chunking significantly outperforms traditional early-chunking baselines, achieving consistent improvements in recall and relevance metrics. Crucially, the method maintains zero-shot deployability—requiring no training or architectural modification—while delivering state-of-the-art retrieval performance.
This study addresses the challenges of text embedding and retrieval for low-resource, morphologically complex agricultural documents in Khmer by systematically evaluating four chunking strategies—recursive, Khmer-aware, sentence-based, and large language model–based—within a retrieval-augmented generation (RAG) framework. Using the BGE-M3 multilingual embedding model and FAISS for dense retrieval, the authors conduct a multidimensional assessment incorporating L2 distance, Khmer IoU, answer relevance, and 5-fold cross-validation. The work reveals, for the first time, the critical impact of chunk granularity and structural preservation on retrieval performance in low-resource languages. The recursive chunking strategy with 300-character segments achieves the best results (L2 distance: 0.4295; answer relevance: 0.8663), significantly outperforming sentence-based chunking (p = 0.0121).
This work addresses the challenge of fragmented evidence in multimodal long-document question answering, where conventional top-k retrieval struggles to model cross-modal associations among text, tables, and slides. The paper introduces a novel formulation that casts evidence assembly as a minimum-cost flow optimization problem over a multimodal node graph. A unified scoring vector jointly governs source/sink selection, edge costs, and capacities, while integrating MMR-based source selection, length-aware answerability proxies, entropy-regularized replicator dynamics, and a dual-process gating mechanism. This enables end-to-end unification of retrieval, routing, selection, and adaptive computation. Evaluated on VisDoMBench, the method substantially outperforms existing baselines, achieving state-of-the-art results on the PaperTab (58.40) and SlideVQA (72.93) subsets, with a macro-average score of 65.47—approaching the strongest baseline, G²-Reader.
This work addresses the limitations of conventional text chunking methods in retrieval-augmented generation (RAG), which often compromise semantic integrity and suffer from a “measurement trap”: removing heading-chain prefixes drastically reduces annotator agreement, casting doubt on the validity of prevailing ablation-based evaluations. To overcome these issues, the authors propose a three-stage, large language model–free semantic chunking pipeline—comprising heading segmentation, semantic merging, and heading-chain prefixing—that leverages inherent document structure to construct contextually coherent chunks and enhance retrieval relevance. Evaluated on a production-scale Markdown knowledge base with 1,600 queries, the approach achieves a 23.8% relative improvement in MRR@5 (from 0.374 to 0.463) overall and an 11.7% gain (from 0.828 to 0.925) on the answerable subset, with inter-annotator agreement reaching Cohen’s κ = 0.45.
This work addresses the limitations of conventional top-k embedding-based retrieval methods when applied to structured long documents such as financial reports, where chunking often severs critical contextual relationships—such as those between numerical values, units, and fiscal year headings—leading to loss of essential information. To overcome this, the authors propose READ, a novel framework that abandons the embedding-retrieval paradigm entirely and instead introduces a deterministic, replayable agent-driven mechanism. READ performs normalized lexical search, structural navigation, and bounded-span reading directly on the original document to retrieve information and generate traceable audit trails. Evaluated on 51 verification questions, READ achieves an accuracy of 58.8%, significantly outperforming dense retrieval (15.7%, p = 2×10⁻⁵), thereby demonstrating the efficacy of embedding-free approaches for precise and interpretable information extraction from structured documents.