A Graph-Based Test-Harness for LLM Evaluation

📅 2025-08-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key limitations of existing LLM evaluation benchmarks for clinical practice guidelines—namely narrow coverage, static design, and heavy reliance on manual annotation—by proposing the first dynamic, systematic framework for assessing LLMs’ guideline-following capabilities. Methodologically, we model the WHO IMCI manual as a directed knowledge graph and automatically generate age-specific, clinically grounded multiple-choice questions via graph traversal, yielding >400 questions and 3.3 trillion answer combinations; contextually plausible distractors further increase difficulty. Our contributions are threefold: (1) enabling automatic benchmark expansion driven by guideline updates; (2) uncovering critical performance gaps—particularly in severity triage, treatment decision-making, and follow-up recommendation—with symptom identification accuracy ranging only from 45% to 67%, and other tasks substantially lower; and (3) generating high-reward, annotation-free samples that effectively support supervised fine-tuning, GRPO, and DPO, ensuring strong scalability and resistance to data contamination.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsSearch and Optimization: Sampling/Simulation-based Search

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
We present a first known prototype of a dynamic, systematic benchmark of medical guidelines for 400+ questions, with 3.3+ trillion possible combinations, covering 100% of guideline relationships. We transformed the WHO IMCI handbook into a directed graph with 200+ nodes (conditions, symptoms, treatments, follow-ups, severities) and 300+ edges, then used graph traversal to generate questions that incorporated age-specific scenarios and contextual distractors to ensure clinical relevance. Our graph-based approach enables systematic evaluation across clinical tasks (45-67% accuracy), and we find models excel at symptom recognition but struggle with triaging severity, treatment protocols and follow-up care, demonstrating how customized benchmarks can identify specific capability gaps that general-domain evaluations miss. Beyond evaluation, this dynamic MCQA methodology enhances LLM post-training (supervised finetuning, GRPO, DPO), where correct answers provide high-reward samples without expensive human annotation. The graph-based approach successfully addresses the coverage limitations of manually curated benchmarks. This methodology is a step toward scalable, contamination-resistant solution for creating comprehensive benchmarks that can be dynamically generated, including when the guidelines are updated. Code and datasets are available at https://github.com/jessicalundin/graph_testing_harness
Problem

Research questions and friction points this paper is trying to address.

Systematically evaluates LLMs on medical guideline comprehension
Identifies specific clinical capability gaps in symptom and treatment tasks
Addresses coverage limitations of manually curated medical benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-based benchmark for systematic medical evaluation
Dynamic question generation via graph traversal method
Automated high-reward samples for LLM post-training
J
Jessica Lundin
Institute for Disease Modeling, Gates Foundation
G
Guillaume Chabot-Couture
Institute for Disease Modeling, Gates Foundation