🤖 AI Summary
This study addresses the limitations of conventional safety metrics in evaluating clinical triage large language models (LLMs) deployed in low-income settings, where apparent compliance may mask systemic unreliability. Leveraging real-world data from 19 primary care clinics in Nigeria, the authors construct IyawoBench v2.0—a benchmark comprising 200 synthetic cases—and introduce the first formal triage safety framework, which decomposes safety into three failure modes and includes 14 formal definitions and two theorems. They propose novel evaluation metrics, including the Escalation Bias Index and Expected Deployment Cost, revealing that all state-of-the-art models exhibit systemic failures. Notably, traditional sensitivity metrics obscure a 77-percentage-point undertriage gap in Llama 3.1 8B. The optimal model choice is shown to be highly dependent on deployment objectives—such as prioritizing emergency care, system sustainability, or balanced performance.
📝 Abstract
Large language models are being deployed as clinical triage tools in low and middle income countries where trained physicians are scarce. Existing safety metrics, however, produce misleading confidence: models scoring 100% on binary "did not send an emergency home" safety measures may nevertheless exhibit systematic failure modes that render them undeployable at scale. We present IyawoBench v2.0, an extended diagnostic evaluation of large language model clinical triage on 200 synthetic vignettes derived from 1,200 real patient encounters at 19 Nigerian primary health centres. We introduce a formal mathematical framework comprising fourteen definitions and two theorems that decompose triage safety into three distinct failure modes: Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability. We propose the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conventional accuracy and sensitivity scores. Evaluated on three frontier models (Claude Sonnet 4.6, Llama 3.3 70B, Llama 3.1 8B) plus five naive baselines, we show that: (1) all three models exhibit at least one formal failure mode; (2) traditional sensitivity metrics conceal a 77 percentage point under-triage gap in Llama 3.1 8B; (3) the optimal model varies across three deployment scenarios (Emergency-Focused, System-Sustainability, Balanced), demonstrating that single-ranking benchmarks are inadequate for LMIC clinical AI selection. IyawoBench v2.0 provides both a rigorous benchmark and a diagnostic framework transferable to any triage-style clinical AI evaluation. All code, data, and analysis pipelines are publicly available.