Large Language Models with Human-In-The-Loop Validation for Systematic Review Data Extraction

πŸ“… 2025-01-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Manual data extraction in systematic reviews is time-consuming, labor-intensive, and heavily reliant on domain expertise, necessitating efficient and reliable automation. Method: We propose AIDE, an AI-augmented human-in-the-loop (HIL) framework leveraging open-source large language models (LLMs)β€”including Gemini 1.5 Flash/Pro and Mistral Large 2β€”to extract 24 key variables from 112 studies. AIDE introduces the first open-source, user-friendly interactive tool that seamlessly integrates LLM inference with human verification, enabling full traceability and end-to-end controllability. Results: Across three LLMs, extraction accuracy reaches 71.17%, 72.14%, and 62.43%, respectively. AIDE supports zero-cost, fully reproducible deployment and significantly improves both extraction efficiency and result credibility. It represents the first lightweight, highly transparent HIL solution for evidence-based research data extraction.

Technology Category

Humans and AI: Human-in-the-loop Machine LearningNatural Language Processing: Information ExtractionData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
πŸ“ Abstract
Systematic reviews are time-consuming endeavors. Historically speaking, knowledgeable humans have had to screen and extract data from studies before it can be analyzed. However, large language models (LLMs) hold promise to greatly accelerate this process. After a pilot study which showed great promise, we investigated the use of freely available LLMs for extracting data for systematic reviews. Using three different LLMs, we extracted 24 types of data, 9 explicitly stated variables and 15 derived categorical variables, from 112 studies that were included in a published scoping review. Overall we found that Gemini 1.5 Flash, Gemini 1.5 Pro, and Mistral Large 2 performed reasonably well, with 71.17%, 72.14%, and 62.43% of data extracted being consistent with human coding, respectively. While promising, these results highlight the dire need for a human-in-the-loop (HIL) process for AI-assisted data extraction. As a result, we present a free, open-source program we developed (AIDE) to facilitate user-friendly, HIL data extraction with LLMs.
Problem

Research questions and friction points this paper is trying to address.

Systematic Reviews
Data Extraction
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large-scale Language Models
Systematic Review Acceleration
Human-AI Collaboration
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Noah L. Schroeder
Noah L. Schroeder
University of Florida
Pedagogical agentsmultimedia learningmeta-analysisvirtual humanssocially interactive agents
C
Chris Davis Jaldi
Department of Computer Science & Engineering, Wright State University, Dayton, Ohio
S
Shan Zhang
School of Teaching and Learning, University of Florida, Gainesville, Florida