Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing literature contains fragmented evidence linking microbes to carcinogenesis, which is challenging for humans to synthesize comprehensively, thereby impeding the discovery of novel oncogenic microbes. This study constructs a structured template and an expert-annotated dataset to systematically evaluate the performance of large language models—including Gemini 2.5 Pro/Flash, GPT-5, and GPT-5 Nano—in tasks involving evidence extraction and critical appraisal. It further introduces a novel evaluation paradigm that quantifies alignment between model outputs and expert judgments. Results demonstrate that the GPT-5 series exhibits no statistically significant differences from expert ratings across multiple question formats (multiple-choice, Likert scale, multiple-select, and free-text responses) and rarely generates hallucinations, providing the first evidence that these models achieve expert-level performance in structured scientific evaluation tasks. However, limitations remain in full-text methodological assessment and identification of contradictory evidence.
📝 Abstract
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement distributions to determine whether LLMs behaved as additional experts by increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Across all question types, LLM responses aligned closely with experts, with GPT-5 and GPT-5 Nano achieving score distributions indistinguishable from experts. Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria. Hallucinations were rare. Methodological appraisal and identification of contradictions within full-texts were the most persistent LLM vulnerabilities. GPT-5 and GPT-5 Nano were indistinguishable from experts on structured domain research paper evaluation tasks. This supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-texts remain weaknesses requiring strengthening.
Problem

Research questions and friction points this paper is trying to address.

microbial oncogenesis
evidence synthesis
systematic review
cancer burden
expert appraisal
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
evidence synthesis
microbial oncogenesis
expert-level evaluation
structured appraisal
🔎 Similar Papers
No similar papers found.
K
Kaela Kokkas
Department of Clinical Microbiology and Infectious Diseases, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
H
Hairong Wang
School of Computer Science and Applied Mathematics, University of the Witwatersrand, Johannesburg, South Africa; Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; Wits Machine Intelligence and Neural Discovery (MIND) Institute, University of the Witwatersrand, Johannesburg, South Africa
Richard Klein
Richard Klein
Associate Professor, School of Computer Science and Applied Mathematics, University of the
Computer VisionDeep LearningMachine LearningArtificial Intelligence
N
Nazir A. Ismail
Department of Clinical Microbiology and Infectious Diseases, National Health Laboratory Service and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
N
Natalie Irwin
Division of Medical Oncology, Department of Internal Medicine, University of the Witwatersrand, Johannesburg, South Africa; Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
M
Mohammad Z. Moonsamy
School of Computer Science and Applied Mathematics, University of the Witwatersrand, Johannesburg, South Africa; Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
K
Kubendran Naidoo
Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; South African Medical Research Council Wits Antiviral Gene Therapy Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; Natio
J
Jeremy Nel
Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; Division of Infectious Diseases, School of Clinical Medicine, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
E
Ekene E. Nweke
Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; Department of Surgery, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
R
Raveen Parboosing
Division of Virology, University of the Witwatersrand & National Health Laboratory Service, Johannesburg, South Africa
E
Emmanuel K. Sekyi
OncoVectra, London, United Kingdom
R
Rebecca T. van Dorsten
Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; South African Medical Research Council Wits Antiviral Gene Therapy Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa
B
Bruce A. Bassett
School of Computer Science and Applied Mathematics, University of the Witwatersrand, Johannesburg, South Africa; Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; Wits Machine Intelligence and Neural Discovery (MIND) Institute, University of the Witwatersrand, Johannesburg, South Africa
R
Robert F. Breiman
Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa; Wits Machine Intelligence and Neural Discovery (MIND) Institute, University of the Witwatersrand, Johannesburg, South Africa; Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA United States