🤖 AI Summary
Existing literature contains fragmented evidence linking microbes to carcinogenesis, which is challenging for humans to synthesize comprehensively, thereby impeding the discovery of novel oncogenic microbes. This study constructs a structured template and an expert-annotated dataset to systematically evaluate the performance of large language models—including Gemini 2.5 Pro/Flash, GPT-5, and GPT-5 Nano—in tasks involving evidence extraction and critical appraisal. It further introduces a novel evaluation paradigm that quantifies alignment between model outputs and expert judgments. Results demonstrate that the GPT-5 series exhibits no statistically significant differences from expert ratings across multiple question formats (multiple-choice, Likert scale, multiple-select, and free-text responses) and rarely generates hallucinations, providing the first evidence that these models achieve expert-level performance in structured scientific evaluation tasks. However, limitations remain in full-text methodological assessment and identification of contradictory evidence.
📝 Abstract
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement distributions to determine whether LLMs behaved as additional experts by increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Across all question types, LLM responses aligned closely with experts, with GPT-5 and GPT-5 Nano achieving score distributions indistinguishable from experts. Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria. Hallucinations were rare. Methodological appraisal and identification of contradictions within full-texts were the most persistent LLM vulnerabilities. GPT-5 and GPT-5 Nano were indistinguishable from experts on structured domain research paper evaluation tasks. This supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-texts remain weaknesses requiring strengthening.