Score
Design and implement methods to segment sequences (documents, transcripts, or temporal streams) into retrieval-ready chunks, assign timestamps/timecodes, encode compact chunk representations, and build indexes and retrieval pipelines that return time-aligned excerpt candidates and ranked chunk lists. Develop and run chunk-level and decoupled retrieval evaluations and benchmarks — including evidence-grounded and timestamped retrieval metrics — to measure chunk retrieval performance, topical relevance, temporal alignment, and to provide candidate pools for downstream selection.
This study addresses the lack of systematic evaluation and inconsistent benchmarks in existing document chunking strategies for dense retrieval. The authors propose the first two-dimensional taxonomy that encompasses structural, semantic-aware, and large language model (LLM)-guided chunking approaches, along with embedding timing considerations. They establish a unified reproducible framework to comprehensively evaluate diverse strategies—including fixed-length, paragraph-level, LumberChunker, and Late Chunking—across both within-document and corpus-level retrieval tasks. Their findings reveal that structural chunking outperforms LLM-based methods in corpus retrieval, while LumberChunker achieves the best performance in within-document retrieval. Notably, contextualized chunking improves corpus retrieval effectiveness but degrades within-document performance, highlighting a task-dependent trade-off that informs optimal chunking selection.
Long-document retrieval faces core challenges including excessive length, scattered evidence, and complex structural organization. This paper systematically surveys the evolution of long-document retrieval techniques across the pre-trained model and large language model (LLM) eras, unifying three developmental stages—classical passage retrieval, hierarchical encoding with efficient attention mechanisms, and LLM-driven re-ranking—for the first time. We propose a unified technical taxonomy encompassing key paradigms such as passage aggregation, structure-aware encoding, and retrieval-augmented generation, alongside domain-specific evaluation resources. Furthermore, we identify critical open problems in the foundation model era: the efficiency–effectiveness trade-off, cross-modal semantic alignment, and result interpretability. This work establishes an authoritative, structured survey framework and a comprehensive roadmap for long-document information retrieval research.
This work addresses the challenge of users struggling to pinpoint specific moments in meeting discussions based solely on content. To overcome this, the paper proposes a novel approach that reframes timestamp prediction as a constrained candidate selection task. Instead of directly generating timestamps, large language models such as Mistral-7B-Instruct are guided to select the most relevant segment from a set of retrieved, timestamped meeting excerpts, thereby avoiding unsupported or invalid predictions. Integrating retrieval-augmented generation (RAG) with a constrained selection mechanism, the method demonstrates significant improvements on a dataset of 200 municipal meetings and 420 queries: Recall@5 increases from 31.9% to 50.0%, mean absolute error decreases to 761 seconds, and the number of valid outputs rises from 373 to 419, substantially enhancing both accuracy and reliability in temporal localization.
Traditional retrieval methods (e.g., BM25, DPR) struggle to jointly optimize topical relevance and temporal alignment in time-sensitive retrieval. To address this, this paper introduces explicit temporal signal modeling into dense passage retrieval for the first time: it jointly encodes query timestamps and document publication dates into the dense representation space of a BERT dual-encoder architecture, and proposes a temporal-aware negative sampling strategy coupled with a contrastive learning objective to enhance temporal semantic discrimination. The core innovations are an end-to-end fusion mechanism for temporal embeddings and a temporally aware training paradigm. Experiments on ArchivalQA and ChroniclingAmericaQA demonstrate significant improvements: Top-1 accuracy increases by 6.63% and 9.56%, respectively, while NDCG@10 improves by 3.79% and 4.68%, substantially outperforming baseline methods.
Existing dense retrieval models underperform on temporally constrained queries (e.g., containing numerical temporal expressions like “in 2015”), while dedicated temporal retrieval approaches often degrade performance on non-temporal queries. To address this trade-off, this paper proposes a temporal identifier-aware model fusion framework. It trains multiple specialized retrievers—each targeting distinct temporal identifiers (e.g., years, seasons, relative temporal terms)—and integrates them via a parameter-efficient fusion mechanism that jointly models temporal and non-temporal semantics without catastrophic forgetting. Experiments across multiple standard benchmarks demonstrate that our method significantly improves retrieval effectiveness for time-sensitive queries (+12.3% average MRR@10), while maintaining or slightly surpassing baseline performance on general queries. This achieves a synergistic optimization of temporal awareness and generalization capability.
To address the degradation in retrieval performance caused by contextual information loss in conventional text chunking for embedding, this paper proposes *late chunking*: chunking is performed after the top-layer Transformer output but before mean pooling, followed by token-level context-aware aggregation to generate chunk embeddings. This approach delays the chunking operation until after global contextual representations have been fully modeled, thereby inherently preserving long-range context without requiring model retraining—enabling plug-and-play adaptation to diverse long-context embedding models. Additionally, a lightweight fine-tuning strategy is introduced to enhance chunk-level discriminability. Extensive evaluation across multiple retrieval benchmarks demonstrates that late chunking significantly outperforms traditional early-chunking baselines, achieving consistent improvements in recall and relevance metrics. Crucially, the method maintains zero-shot deployability—requiring no training or architectural modification—while delivering state-of-the-art retrieval performance.
This work addresses the limitations of traditional Retrieval-Augmented Generation (RAG) systems, which rely on static semantic chunking and struggle to balance retrieval precision and recall while lacking adaptability to query-specific context. The authors propose Query-Adaptive Semantic Chunking (QASC), a novel approach that integrates user query information directly into the chunking phase for the first time. QASC identifies seed sentences based on cosine similarity between sentences and the query, dynamically expands contextual windows, and leverages an embedding model combined with a chunk-level relevance aggregation mechanism to produce semantically coherent and query-relevant text chunks. Evaluated on 100 technical documents and 200 queries, QASC achieves an F1 score of 0.85—outperforming fixed-size chunking by 18–27% and surpassing existing semantic and agent-based chunking methods by 8–12%. High inter-annotator agreement in human evaluation further validates its effectiveness (Cohen’s κ = 0.82).
This work addresses the challenge of large index sizes in multi-vector retrieval models, which stem from their long embedding sequences and hinder practical deployment. The study presents the first systematic evaluation of training-free token compression strategies that directly reduce the sequence dimensionality of multi-vector embeddings to lower memory overhead and query latency. By comparing token merging against token pruning, the authors demonstrate that merging achieves a superior trade-off: it substantially shrinks index size while more effectively preserving retrieval performance. These findings establish token merging as a practical and effective solution for enabling efficient multi-vector retrieval without compromising accuracy.
This work addresses the limitations of conventional text chunking methods in retrieval-augmented generation (RAG), which often compromise semantic integrity and suffer from a “measurement trap”: removing heading-chain prefixes drastically reduces annotator agreement, casting doubt on the validity of prevailing ablation-based evaluations. To overcome these issues, the authors propose a three-stage, large language model–free semantic chunking pipeline—comprising heading segmentation, semantic merging, and heading-chain prefixing—that leverages inherent document structure to construct contextually coherent chunks and enhance retrieval relevance. Evaluated on a production-scale Markdown knowledge base with 1,600 queries, the approach achieves a 23.8% relative improvement in MRR@5 (from 0.374 to 0.463) overall and an 11.7% gain (from 0.828 to 0.925) on the answerable subset, with inter-annotator agreement reaching Cohen’s κ = 0.45.
This study addresses the challenge of temporal inconsistency in semantic retrieval from classical Chinese chronicles, where time expressions are typically implicit and non-Gregorian, often leading to erroneous results. To tackle this issue, the work introduces the first month-granularity time-keyed retrieval task tailored to the *Spring and Autumn Annals* (*Chunqiu*), along with a new benchmark dataset, ChunQiuTR, which includes temporally proximate distractors to rigorously evaluate temporal fidelity. The authors propose a calendar-aware dual-encoder model (CTD) that integrates absolute calendrical information via Fourier encoding and relative temporal offsets. Experimental results demonstrate that the proposed approach significantly outperforms strong semantic dual-encoder baselines on time-keyed retrieval, underscoring the critical role of temporal consistency in retrieval-augmented generation for historical texts.
This work addresses the inefficiencies of conventional text chunking in standard retrieval-augmented generation (RAG) systems, which often introduces redundancy, leading to excessive storage costs and degraded retrieval performance. To mitigate this, the authors propose a lightweight pre-index filtering mechanism that integrates semantic similarity, topic coherence, and named entity recognition to selectively prune redundant text chunks prior to indexing. Evaluated through a token-level precision, recall, and intersection-over-union framework, the approach reduces vector index size by 25%–36% while preserving retrieval quality comparable to that of the original system, thereby significantly enhancing the overall efficiency of RAG pipelines.