🤖 AI Summary
This study addresses the vulnerability of open-source large language models to counterfactual biases driven by non-clinical factors in pediatric emergency triage, which threatens decision fairness. To investigate this, we construct paired counterfactual cases altering only demographic or social variables to audit the sensitivity of Emergency Severity Index predictions across ten open-source models. We propose a lightweight, interpretable framework incorporating hierarchical correlation analysis, evaluated alongside QLoRA fine-tuning and medical-specific models such as MedGemma. Our findings reveal that scaling model size or employing medical pretraining does not necessarily mitigate bias, uncovering latent directional failure modes. Notably, the fine-tuned Qwen2.5-7B achieves the lowest bias rate at 5.27%. This work provides empirical evidence and methodological support for conducting fairness audits prior to clinical deployment.
📝 Abstract
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.