PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

📅 2026-09-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对AI辅助系统评价报告不一致的问题,通过分析SciLitBench数据集,提出PRISMA-LLM框架以改进方法、评估和限制报告。
📝 Abstract
Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Problem

Research questions and friction points this paper is trying to address.

Large language models
AI-assisted systematic reviews
reporting inconsistency
workflow audit
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

PRISMA-LLM
systematic reviews
large language models
reporting framework
workflow transparency