Can Machine Learning Support the Selection of Studies for Systematic Literature Review Updates?

📅 2025-02-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Systematic Literature Review (SLR) updates face a critical trade-off between reducing human effort in study screening and preserving evidence completeness. Method: We propose and empirically evaluate a human-in-the-loop screening framework for SLR updates in software engineering, employing Random Forest and SVM models optimized for 100% recall during pre-screening. Contribution/Results: Our approach reduces manual screening effort by 33.9% while achieving an F1-score of 0.33—confirming its unsuitability as a fully automated replacement but strong value as a high-recall pre-screening tool. Dual-reviewer initial screening yielded results closest to final inclusion decisions. This work presents the first rigorous, recall-guaranteed application of machine learning in SLR updating, coupled with quantitative evaluation of human-AI collaboration efficiency. It provides a reproducible, generalizable methodological foundation for evidence-driven automation of SLRs.

Technology Category

Humans and AI: Human-in-the-loop Machine LearningMachine Learning: Learning Preferences or RankingsSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
[Background] Systematic literature reviews (SLRs) are essential for synthesizing evidence in Software Engineering (SE), but keeping them up-to-date requires substantial effort. Study selection, one of the most labor-intensive steps, involves reviewing numerous studies and requires multiple reviewers to minimize bias and avoid loss of evidence. [Objective] This study aims to evaluate if Machine Learning (ML) text classification models can support reviewers in the study selection for SLR updates. [Method] We reproduce the study selection of an SLR update performed by three SE researchers. We trained two supervised ML models (Random Forest and Support Vector Machines) with different configurations using data from the original SLR. We calculated the study selection effectiveness of the ML models for the SLR update in terms of precision, recall, and F-measure. We also compared the performance of human-ML pairs with human-only pairs when selecting studies. [Results] The ML models achieved a modest F-score of 0.33, which is insufficient for reliable automation. However, we found that such models can reduce the study selection effort by 33.9% without loss of evidence (keeping a 100% recall). Our analysis also showed that the initial screening by pairs of human reviewers produces results that are much better aligned with the final SLR update result. [Conclusion] Based on our results, we conclude that although ML models can help reduce the effort involved in SLR updates, achieving rigorous and reliable outcomes still requires the expertise of experienced human reviewers for the initial screening phase.
Problem

Research questions and friction points this paper is trying to address.

Machine Learning aids study selection
Reduces effort in SLR updates
Human expertise remains crucial
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine Learning text classification
Random Forest and SVM models
Human-ML pair study selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marcelo Costalonga
PUC-Rio, Rio de Janeiro, Brazil
B
B. Napoleão
Université du Québec à Chicoutimi, Chicoutimi, Canada
M
M. T. Baldassarre
University of Bari, Bari, Italy
K
K. Felizardo
Universidade Tecnológica Federal do Paraná, Cornélio Procópio, Brazil
Igor Steinmacher
Igor Steinmacher
Northern Arizona University
Software EngineeringCSCWMining Software RepositoriesOpen Source Software
Marcos Kalinowski
Marcos Kalinowski
Professor, Pontifical Catholic University of Rio de Janeiro (PUC-Rio)
Empirical Software EngineeringAI EngineeringAI4SEHuman Aspects in Software Engineering