An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques

📅 2025-07-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the zero-shot summarization capabilities of six large language models (LLMs) across diverse domains—including news, dialogue, and scientific literature—with emphasis on few-shot constraints and long-document challenges. To address context-length limitations, we propose a sentence-based chunking strategy enabling short-context models to process lengthy scientific papers in stages, thereby improving summary quality. Our methodology integrates zero-shot prompting, in-context learning, and dual-metric evaluation using ROUGE and BERTScore, alongside inference efficiency analysis. Results show that LLMs excel in news and dialogue summarization; the chunking strategy boosts average ROUGE-L scores for scientific literature by 12.3%; and strong interaction effects emerge among model scale, domain specificity, and prompt design. This work provides a reproducible methodology and empirical benchmark for lightweight, instruction-driven summarization systems.

Technology Category

Natural Language Processing: SummarizationMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Data Visualization & Summarization

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Large Language Models (LLMs) continue to advance natural language processing with their ability to generate human-like text across a range of tasks. Despite the remarkable success of LLMs in Natural Language Processing (NLP), their performance in text summarization across various domains and datasets has not been comprehensively evaluated. At the same time, the ability to summarize text effectively without relying on extensive training data has become a crucial bottleneck. To address these issues, we present a systematic evaluation of six LLMs across four datasets: CNN/Daily Mail and NewsRoom (news), SAMSum (dialog), and ArXiv (scientific). By leveraging prompt engineering techniques including zero-shot and in-context learning, our study evaluates the performance using the ROUGE and BERTScore metrics. In addition, a detailed analysis of inference times is conducted to better understand the trade-off between summarization quality and computational efficiency. For Long documents, introduce a sentence-based chunking strategy that enables LLMs with shorter context windows to summarize extended inputs in multiple stages. The findings reveal that while LLMs perform competitively on news and dialog tasks, their performance on long scientific documents improves significantly when aided by chunking strategies. In addition, notable performance variations were observed based on model parameters, dataset properties, and prompt design. These results offer actionable insights into how different LLMs behave across task types, contributing to ongoing research in efficient, instruction-based NLP systems.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLMs' text summarization across diverse domains and datasets
Assessing summarization without extensive training data as a bottleneck
Exploring prompt engineering impact on LLM performance and efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Utilizes prompt engineering techniques for evaluation
Implements sentence-based chunking for long documents
Evaluates models using ROUGE and BERTScore metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Walid Mohamed Aly
Information Systems Department, Faculty of Computers and Information, Assiut University, Assiut, Egypt, 71515
Taysir Hassan A. Soliman
Taysir Hassan A. Soliman
Dean of Faculty of Computers and Information, Assuit University, Egypt
bioinformaticsBig data analyticsData Scienceand data mining
A
Amr Mohamed AbdelAziz
Information Systems Department, Faculty of Computers and Artificial Intelligence, Beni-Suef University, Beni-Suef, Egypt, 62111