A Benchmark Framework for Screening Automation in Systematic Reviews

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the time-intensive nature of literature screening in systematic reviews and the failure of conventional evaluation metrics under class imbalance. To overcome these challenges, we construct a benchmark dataset comprising 45,000 records and introduce the first evaluation framework specifically designed for highly imbalanced data. Furthermore, we develop PromptSR, a tool that facilitates automated screening using large language models (LLMs). By integrating prompt engineering with techniques for handling imbalanced data, this work provides a comprehensive toolchain spanning from standardized benchmarks to experiment management. Ultimately, this research effectively resolves critical evaluation bottlenecks in LLM-assisted screening, significantly enhancing both screening efficiency and evaluation accuracy.
📝 Abstract
Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets. This paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces PromptSR, a tool designed to support prompt experimentation, experiment management, and result analysis for LLM-based screening. We also present a use case demonstrating the application of SRBench and PromptSR.
Problem

Research questions and friction points this paper is trying to address.

Systematic Reviews
Screening Automation
Large Language Models
Class Imbalance
Benchmark Framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Systematic Review Screening
Benchmark Framework
Class Imbalance
Prompt Engineering
💼 Related Jobs
No related jobs found.
G
Gauransh Kumar
Université de Montréal, Montreal, Canada
Luciano Marchezan
Luciano Marchezan
Université de Montréal, Montreal, Canada
G
Guillaume Genois
Université de Montréal, Montreal, Canada
K
Kévin Delcourt
Université de Montréal, Montreal, Canada
Eugene Syriani
Eugene Syriani
Université de Montréal
Software engineeringmodel-driven engineeringmodel transformationsimulation