Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of open-source large language models to counterfactual biases driven by non-clinical factors in pediatric emergency triage, which threatens decision fairness. To investigate this, we construct paired counterfactual cases altering only demographic or social variables to audit the sensitivity of Emergency Severity Index predictions across ten open-source models. We propose a lightweight, interpretable framework incorporating hierarchical correlation analysis, evaluated alongside QLoRA fine-tuning and medical-specific models such as MedGemma. Our findings reveal that scaling model size or employing medical pretraining does not necessarily mitigate bias, uncovering latent directional failure modes. Notably, the fine-tuned Qwen2.5-7B achieves the lowest bias rate at 5.27%. This work provides empirical evidence and methodological support for conducting fairness audits prior to clinical deployment.
📝 Abstract
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.
Problem

Research questions and friction points this paper is trying to address.

Counterfactual bias
Clinical triage
Large language models
Fairness auditing
Emergency Severity Index
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Auditing
Large Language Models
Clinical Triage
Bias Mitigation
Fairness
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Manar Aljohani
Department of Computer Science, Virginia Tech
Brandon Ho
Brandon Ho
Ph.D. Candidate Robotics, Georgia Institute of Technology
Controls TheoryArtificial IntelligenceMechanics
K
Kenneth McKinley
Children’s National Hospital
D
Dennis Ren
Children’s National Hospital
X
Xuan Wang
Department of Computer Science, Virginia Tech