Score
Designs, implements, and evaluates models that produce concise representations of longer inputs by selecting salient content and/or generating abstractive rewrites, covering extractive, abstractive, hybrid, hierarchical multi-document, two-stage extract-then-abstract, and sequence-level summarization approaches for text and code. Builds and analyzes components for passage/sentence salience and importance estimation, highlight/crowd-salience prediction, and ranking (e.g., precision@k) along with appropriate model architectures, training procedures, and evaluation pipelines.
Large language models (LLMs) exhibit unstable summarization quality and limited controllability over abstraction levels. Method: This paper proposes a controllable abstractive summarization framework based on multi-stage prompt engineering, integrating semantic analysis, topic modeling, and noise-aware control to enable adjustable abstraction granularity. We systematically investigate the impact of prompt length, data noise, and text genre on summarization performance using the CNN/Daily Mail benchmark. Contribution/Results: Experiments demonstrate that medium-length prompts yield statistically significant improvements in ROUGE-L scores; increased input noise degrades performance consistently; and LLMs generalize best on news-domain texts. The framework provides an interpretable, configurable pathway to enhance accuracy, consistency, and abstraction-level control in LLM-generated summaries.
Existing summarization evaluation metrics, such as ROUGE, struggle to effectively quantify the abstractiveness of generated summaries and fail to capture the fundamental distinction between extractive and abstractive approaches. This work proposes a novel framework for measuring abstractiveness through three heuristic indicators: Reference Abstractiveness (RA), Summary Abstractiveness (SA), and Abstractiveness Ratio (AR), augmented by document-length modulation and a cubic non-overlap factor to quantify the degree of deviation from the source text. Notably, the introduction of AR enables the detection of potential hallucinations, yielding an evaluation system that is dimensionally consistent, bounded, and nonlinearly sensitive to the extractive–abstractive boundary. Experiments on XSUM demonstrate the metric’s efficacy, clearly differentiating extractive models (SA ≈ 0.12–0.26) from abstractive ones (SA ≈ 0.96–1.77), thereby validating its practical utility and theoretical soundness.
This work investigates the implicit information salience mechanisms underlying large language models’ (LLMs) text summarization behavior. Addressing the opacity of LLMs’ internal salience preferences, we propose the first behaviorally interpretable salience quantification framework: it jointly models length-controllable summarization with Questions Under Discussion (QUD) answerability tracking to systematically uncover models’ information selection biases. Empirical evaluation across 13 LLMs and 4 benchmark datasets reveals a consistent, architecture- and scale-invariant hierarchical salience pattern—one that is introspectively inaccessible to the models themselves, only weakly aligned with human intuition, and not directly recoverable from internal signals such as attention weights or gradients. Our study establishes a novel paradigm for probing LLM summarization cognition and provides a reproducible, behavior-based analytical toolkit for salience assessment.
To address the challenge in query-focused summarization (QFS) where large language models struggle to jointly model long documents and achieve fine-grained query alignment, this paper proposes the Query-aware HyperExpert framework and the Query-focused Infini-attention mechanism. The former employs query-aware hyper-expert routing for modular semantic adaptation, while the latter integrates query-driven sparse attention constraints into infinite-context modeling. By synergistically combining extractive summarization paradigms with the strong generative capabilities of LLMs, our approach achieves significant improvements over state-of-the-art methods across multiple standard QFS benchmarks. It delivers high accuracy, strong generalization across diverse query types and document domains, and improved inference efficiency. The implementation is publicly available.
To address the error accumulation and high training costs arising from the long-standing separation of extractive and abstractive summarization, this paper proposes the ExtAbs paradigm: an end-to-end joint modeling framework within a unified encoder-decoder architecture. Its core innovation is a parameter-free saliency masking mechanism that dynamically modulates cross-attention weights to explicitly guide the decoder toward salient input segments. By eliminating conventional multi-stage training and auxiliary parameterized modules, ExtAbs enables synergistic optimization of extraction and abstraction. Built upon BART and PEGASUS backbones, ExtAbs achieves state-of-the-art extractive performance across CNN/DailyMail, XSum, and Newsroom benchmarks, while generating summaries on par with—or even surpassing—those of the original large models. These results validate the effectiveness and generalizability of lightweight, unified modeling for dual-purpose summarization.
This study investigates the implicit mechanisms by which large language models (LLMs) assess information importance during abstractive summarization, a process that typically lacks transparency. By generating length-controlled summaries to construct empirical importance distributions and combining attention head alignment analysis with cross-layer predictive capability evaluation, the work systematically reveals consistent and model-family-specific importance patterns in LLMs. The findings demonstrate that LLMs’ importance judgments differ markedly from those of traditional models, that consistency within a model family outweighs the influence of model scale, and that specific attention heads in middle-to-late layers effectively predict salient content. This research is the first to localize neural components aligned with importance judgment, establishing a foundation for interpretable understanding of LLM-based summarization mechanisms.
This study addresses the lack of multidimensional, systematic tools for evaluating text simplification by large language models that simultaneously meet research and educational requirements. To bridge this gap, we propose an interactive human-in-the-loop web application enabling parallel simplification and real-time comparative analysis across diverse prompt–model (P×M) configurations, tailored to any target CEFR proficiency level. The core innovation lies in a visualization mechanism that integrates a hierarchical semantic alignment engine with a linear bias heuristic (λ), substantially reducing cognitive load during manual evaluation and facilitating reproducible, structured annotations. The system combines LLM APIs, semantic alignment algorithms, and a responsive front-end framework. Both source code and a live demonstration platform are publicly released, and the tool is readily applicable to downstream NLP dataset construction.
Existing single-model summarization approaches often exhibit insufficient robustness and inconsistent output quality when handling texts with diverse structures and topics. To address this limitation, this work proposes a multi-model adaptive selection framework that integrates multiple fine-tuned Transformer-based summarization models to generate candidate summaries. The framework employs automatic evaluation metrics such as BERTScore to comprehensively assess candidates at both semantic and lexical levels, enabling adaptive selection of the highest-quality summary. Evaluated on the CNN/DailyMail dataset, the proposed method achieves a BERTScore of 88.63%, significantly outperforming prominent large language models including GPT-3-D2, Falcon-7B, and MPT-7B. This approach effectively enhances both the robustness and generation quality of abstractive summarization systems.
This work aims to enhance the quality and accuracy of abstractive text summarization on the English subset of the XL-Sum corpus. Building upon the Transformer-based PEGASUS model, we fine-tune it on the English XL-Sum data and evaluate its performance using ROUGE metrics. Experimental results demonstrate that our approach substantially outperforms the mT5 baseline, achieving relative improvements of 4.04%, 15.25%, and 3.39% in ROUGE-1, ROUGE-2, and ROUGE-L scores, respectively. To the best of our knowledge, this constitutes the state-of-the-art performance for abstractive summarization on this dataset.
This study addresses the challenge of scientific papers being inaccessible to non-specialist readers due to linguistic complexity. The authors propose a human–AI collaborative simplification pipeline that first leverages GPT-4o-mini to generate initial simplified summaries, followed by iterative refinement through a two-stage feedback loop involving both lay readers and domain experts. This approach innovatively integrates dual human feedback mechanisms to simultaneously enhance readability and preserve terminological accuracy. The project also introduces the first corpus of scientifically simplified texts designed for interdisciplinary communication, annotated with both human judgments and automatic evaluation metrics. Experimental results demonstrate that LLM-generated simplifications are consistently preferred for their clarity and conciseness, while expert editing effectively retains essential technical terms and the strength of scientific claims.