Score
Designs and implements methods to partition text into topically coherent passages by identifying topical boundaries and producing concise, evidence-rich passages. Ensures passage-level normalization of length and granularity and preserves citation and provenance metadata for each passage.
This study addresses the lack of systematic evaluation and inconsistent benchmarks in existing document chunking strategies for dense retrieval. The authors propose the first two-dimensional taxonomy that encompasses structural, semantic-aware, and large language model (LLM)-guided chunking approaches, along with embedding timing considerations. They establish a unified reproducible framework to comprehensively evaluate diverse strategies—including fixed-length, paragraph-level, LumberChunker, and Late Chunking—across both within-document and corpus-level retrieval tasks. Their findings reveal that structural chunking outperforms LLM-based methods in corpus retrieval, while LumberChunker achieves the best performance in within-document retrieval. Notably, contextualized chunking improves corpus retrieval effectiveness but degrades within-document performance, highlighting a task-dependent trade-off that informs optimal chunking selection.
Existing PDF parsing tools are primarily designed for academic papers and struggle to accurately process pedagogical documents—such as legal textbooks—that exhibit complex, implicitly structured hierarchies. To address this, we propose a hierarchical text segmentation framework integrating structure-aware preprocessing with large language models (LLMs). Our method jointly leverages OCR-based heading detection, XML structural feature extraction, and contextual semantic modeling to infer implicit heading hierarchies without requiring explicit table-of-contents input. Compared to pure LLM–based or traditional rule-based approaches, our framework significantly reduces false positives and improves segmentation accuracy. When high-quality metadata is available, a supplementary table-of-contents–driven strategy further enhances performance. The source code and benchmark dataset are publicly released to support reproducible research.
This study addresses the limitations of existing automatic multi-label classification approaches for scholarly papers, which predominantly rely on titles and abstracts—often insufficient for accurate labeling—while full texts are lengthy and exhibit uneven information distribution. To overcome this, the authors propose segmenting full-text articles by physical position and systematically evaluating the discriminative power of individual sections and their combinations. Their analysis reveals that middle-to-late and concluding sections carry higher informational value. By integrating these informative segments with bibliographic metadata, they construct an enhanced multi-label classification model. Experiments on a corpus of 1,954 library and information science journal articles demonstrate that the proposed cross-segment combination and metadata fusion strategy significantly improves classification accuracy, offering a novel and effective pathway for fine-grained method identification in academic texts.
To address insufficient content controllability when compressing lengthy technical documents (e.g., research papers) into concise formats (e.g., blog posts), this paper proposes an interactive reverse-source outline mechanism. It explicitly models LLM-based summarization as an editable, traceable, structured outline—comprising hierarchical, semantically grounded nodes—enabling iterative user refinement of functional points to precisely govern information coverage. The method integrates large language models, dynamic content-to-outline mapping, and an interactive UI to realize an end-to-end, fine-grained summarization system. Empirical evaluation and real-world deployment demonstrate significant improvements: author satisfaction with content coverage increases markedly; information change per editing operation rises by 37%; and retention of critical research insights improves 2.1×. This work achieves, for the first time, bidirectional controllability—both selection and synthesis—over content in technical document summarization.
This work addresses the prevailing limitation in synthetic data generation, which has largely focused on content creation or local rewriting while neglecting the impact of book-level structure on language model training. The authors propose a scalable synthesis pipeline that leverages topic clustering, hierarchical outline planning, and paragraph-aligned generation to construct 686K structured synthetic textbooks comprising 32 billion tokens. Through controlled ablation experiments—the first of their kind—they demonstrate that training on data with coherent book-like organization yields significantly better model performance compared to baselines such as unstructured splitting (Split), random concatenation (RandomConcat), or simple rephrasing (Rephrase). The structured data consistently improves results across multiple downstream tasks, achieving an average gain of +1.09, thereby revealing the critical role of document organization in enhancing training data quality.
This study addresses the longstanding scarcity of large-scale, high-quality annotated data for automatic identification of rhetorical structures—such as Introduction, Methods, Results, and Discussion—in scientific papers. Leveraging the S2ORC corpus, the authors employ a rule-based classification algorithm to automatically annotate section-level rhetorical structures across 15.6 million STEM papers, yielding the first dataset of rhetorical annotations at a ten-million-paper scale. Rigorous evaluation through human assessment and validation by large language models demonstrates that the annotation quality closely aligns with manual labeling. The resulting dataset spans multiple disciplines, including medicine and biology, providing a high-quality resource for large-scale computational analysis of scientific writing patterns.
To address factual inconsistency, information loss, and poor traceability in long-document summarization, this paper proposes a sentence-level highlighting-guided self-planning generation framework. First, it identifies salient sentences via importance modeling and generates a traceable content plan; subsequently, summary generation is conditioned on this plan, effectively decoupling content selection from surface realization. This novel paradigm significantly enhances summary faithfulness and fine-grained detail retention. On the GovReport benchmark, our approach achieves a +4.1-point improvement in ROUGE-L and a 35% gain in SummaC score. Qualitative analysis confirms more complete preservation of critical details, as well as improved cross-domain accuracy and analytical depth in generated summaries.
This study investigates the design of effective text chunking strategies within the framework of the German Civil Code to enhance the performance of Retrieval-Augmented Generation (RAG) systems on legal question-answering tasks. The authors systematically evaluate a range of chunking approaches, including those based on legal structure (articles, paragraphs, sentences, and propositions), fixed-size windows, context-aware segmentation, semantic clustering, and hierarchical retrieval via RAPTOR. Experimental results demonstrate that strategies preserving the inherent legal structure achieve significantly higher recall than more complex semantic methods, while also offering superior efficiency in terms of query latency, index construction time, and storage overhead. These findings highlight a critical trade-off between semantic enrichment and computational cost in legal RAG applications.
Traditional topic models assign a single topic to an entire document, which struggles to accurately represent multi-topic texts and often leads to topic mixing and reduced interpretability. This work proposes a Segment-Based Topic Assignment (SBTA) framework that, for the first time, refines the granularity of topic modeling from the document level to semantically coherent text segments. To support this approach, we construct the SemEval-STM dataset by combining large language model–based automatic segmentation with human refinement to produce high-quality segments, and introduce a segment-level word intrusion task to enable fine-grained evaluation. Experiments demonstrate that SBTA significantly improves topic clustering quality and interpretability across multiple topic models and evaluation metrics, confirming its effectiveness and scalability.
This work addresses the challenge of inconsistent README quality, which stems from varying audiences and usage contexts, and the inability of existing tools to simultaneously accommodate style, content, and contextual appropriateness. The paper proposes LintMe, a novel linter that uniquely integrates programmatic rules with large language model (LLM)-based content understanding. LintMe enables users to define context-sensitive checking rules via a lightweight domain-specific language (DSL), combining programmatic validations—such as link verification—with LLM-driven semantic assessments like terminology recognition. This approach enhances documentation quality while preserving authorial autonomy. A user study (N=11) demonstrates that LintMe is both usable and flexible, significantly outperforming baseline approaches that rely solely on direct LLM usage, and its scalability is further validated through illustrative case studies.