passage segmentation

Designs and implements methods to partition text into topically coherent passages by identifying topical boundaries and producing concise, evidence-rich passages. Ensures passage-level normalization of length and granularity and preserves citation and provenance metadata for each passage.

passagesegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

HiPS: Hierarchical PDF Segmentation of Textbooks

Aug 31, 2025
SW
Sabine Wehnert
🏛️ Otto von Guericke University | Leibniz Institute for Educational Media | Georg Eckert Institute

Existing PDF parsing tools are primarily designed for academic papers and struggle to accurately process pedagogical documents—such as legal textbooks—that exhibit complex, implicitly structured hierarchies. To address this, we propose a hierarchical text segmentation framework integrating structure-aware preprocessing with large language models (LLMs). Our method jointly leverages OCR-based heading detection, XML structural feature extraction, and contextual semantic modeling to infer implicit heading hierarchies without requiring explicit table-of-contents input. Compared to pure LLM–based or traditional rule-based approaches, our framework significantly reduces false positives and improves segmentation accuracy. When high-quality metadata is available, a supplementary table-of-contents–driven strategy further enhances performance. The source code and benchmark dataset are publicly released to support reproducible research.

Evaluating TOC-based and structure-aware methods for textbook analysisHierarchical segmentation of complex structured PDF documentsImproving parsing accuracy with preprocessing and LLM integration

This study addresses the limitations of existing automatic multi-label classification approaches for scholarly papers, which predominantly rely on titles and abstracts—often insufficient for accurate labeling—while full texts are lengthy and exhibit uneven information distribution. To overcome this, the authors propose segmenting full-text articles by physical position and systematically evaluating the discriminative power of individual sections and their combinations. Their analysis reveals that middle-to-late and concluding sections carry higher informational value. By integrating these informative segments with bibliographic metadata, they construct an enhanced multi-label classification model. Experiments on a corpus of 1,954 library and information science journal articles demonstrate that the proposed cross-segment combination and metadata fusion strategy significantly improves classification accuracy, offering a novel and effective pathway for fine-grained method identification in academic texts.

automatic classificationfull-text analysisknowledge extraction

Papers-to-Posts: Supporting Detailed Long-Document Summarization with an Interactive LLM-Powered Source Outline

Jun 14, 2024
MR
Marissa Radensky
🏛️ University of Washington | Allen Institute for AI

To address insufficient content controllability when compressing lengthy technical documents (e.g., research papers) into concise formats (e.g., blog posts), this paper proposes an interactive reverse-source outline mechanism. It explicitly models LLM-based summarization as an editable, traceable, structured outline—comprising hierarchical, semantically grounded nodes—enabling iterative user refinement of functional points to precisely govern information coverage. The method integrates large language models, dynamic content-to-outline mapping, and an interactive UI to realize an end-to-end, fine-grained summarization system. Empirical evaluation and real-world deployment demonstrate significant improvements: author satisfaction with content coverage increases markedly; information change per editing operation rises by 37%; and retention of critical research insights improves 2.1×. This work achieves, for the first time, bidirectional controllability—both selection and synthesis—over content in technical document summarization.

Enabling controlled summarization of long technical documentsFacilitating interactive adjustment of key details in summariesImproving content coverage in detailed long-form summaries

This work addresses the prevailing limitation in synthetic data generation, which has largely focused on content creation or local rewriting while neglecting the impact of book-level structure on language model training. The authors propose a scalable synthesis pipeline that leverages topic clustering, hierarchical outline planning, and paragraph-aligned generation to construct 686K structured synthetic textbooks comprising 32 billion tokens. Through controlled ablation experiments—the first of their kind—they demonstrate that training on data with coherent book-like organization yields significantly better model performance compared to baselines such as unstructured splitting (Split), random concatenation (RandomConcat), or simple rephrasing (Rephrase). The structured data consistently improves results across multiple downstream tasks, achieving an average gain of +1.09, thereby revealing the critical role of document organization in enhancing training data quality.

book-level organizationdocument coherencelanguage model pre-training

Latest Papers

What's happening recently
View more

This study addresses the longstanding scarcity of large-scale, high-quality annotated data for automatic identification of rhetorical structures—such as Introduction, Methods, Results, and Discussion—in scientific papers. Leveraging the S2ORC corpus, the authors employ a rule-based classification algorithm to automatically annotate section-level rhetorical structures across 15.6 million STEM papers, yielding the first dataset of rhetorical annotations at a ten-million-paper scale. Rigorous evaluation through human assessment and validation by large language models demonstrates that the annotation quality closely aligns with manual labeling. The resulting dataset spans multiple disciplines, including medicine and biology, providing a high-quality resource for large-scale computational analysis of scientific writing patterns.

large-scale datasetrhetorical sectionsscientific papers

Enhancing Long Document Long Form Summarisation with Self-Planning

Dec 18, 2025
XD
Xiaotang Du
🏛️ University of Edinburgh

To address factual inconsistency, information loss, and poor traceability in long-document summarization, this paper proposes a sentence-level highlighting-guided self-planning generation framework. First, it identifies salient sentences via importance modeling and generates a traceable content plan; subsequently, summary generation is conditioned on this plan, effectively decoupling content selection from surface realization. This novel paradigm significantly enhances summary faithfulness and fine-grained detail retention. On the GovReport benchmark, our approach achieves a +4.1-point improvement in ROUGE-L and a 35% gain in SummaC score. Qualitative analysis confirms more complete preservation of critical details, as well as improved cross-domain accuracy and analytical depth in generated summaries.

Enhances traceability and faithfulness of generated summaries.Improves factual consistency in long document summarization.Preserves important details for accurate and insightful summaries.

This study investigates the design of effective text chunking strategies within the framework of the German Civil Code to enhance the performance of Retrieval-Augmented Generation (RAG) systems on legal question-answering tasks. The authors systematically evaluate a range of chunking approaches, including those based on legal structure (articles, paragraphs, sentences, and propositions), fixed-size windows, context-aware segmentation, semantic clustering, and hierarchical retrieval via RAPTOR. Experimental results demonstrate that strategies preserving the inherent legal structure achieve significantly higher recall than more complex semantic methods, while also offering superior efficiency in terms of query latency, index construction time, and storage overhead. These findings highlight a critical trade-off between semantic enrichment and computational cost in legal RAG applications.

chunkingGerman statutory lawlegal information retrieval

Traditional topic models assign a single topic to an entire document, which struggles to accurately represent multi-topic texts and often leads to topic mixing and reduced interpretability. This work proposes a Segment-Based Topic Assignment (SBTA) framework that, for the first time, refines the granularity of topic modeling from the document level to semantically coherent text segments. To support this approach, we construct the SemEval-STM dataset by combining large language model–based automatic segmentation with human refinement to produce high-quality segments, and introduce a segment-level word intrusion task to enable fine-grained evaluation. Experiments demonstrate that SBTA significantly improves topic clustering quality and interpretability across multiple topic models and evaluation metrics, confirming its effectiveness and scalability.

document segmentationmulti-theme documentstopic assignment

This work addresses the challenge of inconsistent README quality, which stems from varying audiences and usage contexts, and the inability of existing tools to simultaneously accommodate style, content, and contextual appropriateness. The paper proposes LintMe, a novel linter that uniquely integrates programmatic rules with large language model (LLM)-based content understanding. LintMe enables users to define context-sensitive checking rules via a lightweight domain-specific language (DSL), combining programmatic validations—such as link verification—with LLM-driven semantic assessments like terminology recognition. This approach enhances documentation quality while preserving authorial autonomy. A user study (N=11) demonstrates that LintMe is both usable and flexible, significantly outperforming baseline approaches that rely solely on direct LLM usage, and its scalability is further validated through illustrative case studies.

authorial agencycontext-specific checksdocumentation quality

Hot Scholars

GZ

Guodong Zhou

Soochow University, China
Natural Language ProcessingArtificial Intelligence
JG

Jiafeng Guo

Professor, Institute of Computing Techonology, CAS
Information RetrievalMachine LearningText AnalysisNeuIR
MM

Mandar Mitra

Indian Statistical Institute
Information Retrieval
SS

Sourav Saha

Indian Statistical Institute
Information RetrievalMachine Learning
EN

Ercong Nie

LMU Munich, MCML
Computational LinguisticsNatural Language Processing