MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of auditable evaluation benchmarks for biomedical large language models (LLMs) in multiple sclerosis (MS) MRI research by proposing an end-to-end automated framework that transforms unstructured literature into standardized assessments. By integrating expert indexing, retrieval-augmented generation (RAG), evidence provenance tracking, and automated quality auditing, the framework constructs a multiple-choice benchmark focused on textual reasoning rather than image interpretation. The resulting dataset comprises 3,058 questions spanning 16 subtopics. Evaluation across 12 LLMs reveals a maximum accuracy gap of 42.8%, demonstrating the benchmark's strong discriminative power and its effectiveness in differentiating model capabilities within this specialized domain.
📝 Abstract
Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model Evaluation
Multiple Sclerosis MRI
Benchmark Construction
Biomedical Knowledge
Source-Grounded Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

benchmark construction
source-grounded evaluation
automated quality audit
large language models
multiple sclerosis MRI
🔎 Similar Papers
No similar papers found.