Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of methodological guidance for ensuring rigor, reliability, and transparency in the use of generative artificial intelligence (GenAI) within systematic literature reviews. To bridge this gap, the authors integrate rapid review methodologies, thought experiments, and expert insights to propose GUEST—the first framework specifically designed for the application and evaluation of GenAI in systematic reviews. The framework delineates clear boundaries for appropriate GenAI use, underscores the indispensable role of human oversight, and offers actionable recommendations to enhance the credibility and reproducibility of the review process. By doing so, it provides researchers with a structured approach to responsibly leverage GenAI while maintaining scholarly standards.
📝 Abstract
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.
Problem

Research questions and friction points this paper is trying to address.

Generative AI
Systematic Literature Review
Large Language Models
Research Rigor
Tool Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative AI
Systematic Literature Review
LLM evaluation
GUEST framework
Human-in-the-loop
🔎 Similar Papers
No similar papers found.