Structure-Aware Semantic Chunking with Title-Chain Prefixes: A 1600-Query Evaluation and the Measurement Trap in Text-Transform Ablations

📅 2026-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of conventional text chunking methods in retrieval-augmented generation (RAG), which often compromise semantic integrity and suffer from a “measurement trap”: removing heading-chain prefixes drastically reduces annotator agreement, casting doubt on the validity of prevailing ablation-based evaluations. To overcome these issues, the authors propose a three-stage, large language model–free semantic chunking pipeline—comprising heading segmentation, semantic merging, and heading-chain prefixing—that leverages inherent document structure to construct contextually coherent chunks and enhance retrieval relevance. Evaluated on a production-scale Markdown knowledge base with 1,600 queries, the approach achieves a 23.8% relative improvement in MRR@5 (from 0.374 to 0.463) overall and an 11.7% gain (from 0.828 to 0.925) on the answerable subset, with inter-annotator agreement reaching Cohen’s κ = 0.45.
📝 Abstract
Chunking is the first and most consequential step in retrieval-augmented generation (RAG): every downstream retrieval decision inherits the chunk boundaries. We present a three-stage, chunk-side-only semantic chunking pipeline---header-split, semantic merge, and title-chain prefixing---that costs zero additional LLM calls: the title chain reuses the document's own header hierarchy instead of a generated summary. On a 1600-query stratified evaluation over a production Markdown knowledge base, the pipeline improves MRR@5 from 0.374 to 0.463 (+23.8%) on the full set and from 0.828 to 0.925 (+11.7%) on the answerable subset (n=563), with dual-annotator Cohen's kappa 0.45 (unweighted, 16,000 score pairs). We then report what we tried and what failed: three query-side or architecture-level follow-ups are design dead ends (unevaluated---no comparable run artifacts), one measured failure (prefix weight decay), and one protocol-level failure that is the paper's central methodological finding. In a same-pool prefix on/off ablation, dual-annotator agreement collapsed from kappa 0.45 to 0.04 under identical prompts---stripping the title-chain context strips the disambiguation signal annotators need to agree on relevance. This measurement trap invalidates a common evaluation practice in chunking research and motivates retrieval-time per-candidate prefix evaluation, the direction we recommend from all our evidence.
Problem

Research questions and friction points this paper is trying to address.

chunking
retrieval-augmented generation
evaluation methodology
measurement trap
semantic chunking
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic chunking
title-chain prefixing
structure-aware retrieval
measurement trap
RAG
🔎 Similar Papers
No similar papers found.