Few-Shot Large Language Models for Actionable Triage Categorization of Online Patient Inquiries

📅 2026-05-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of accurately triaging informal and often incomplete online patient queries into four actionable categories—self-care, scheduled appointment, urgent review, and emergency referral—under conditions of scarce labeled data. It presents the first real-world evaluation of few-shot large language models (LLMs) for clinical triage, combining human-calibrated and automatically annotated data to assess six LLMs alongside TF-IDF and BioBERT baselines across 0-shot to 12-shot settings. Results show that Claude Haiku 4.5 with 12 examples achieves a macro-F1 score of 0.475, significantly outperforming the best supervised baseline (BioBERT, 0.378). The work further introduces safety-aware metrics and dual-model consistency analysis, demonstrating that LLMs can support triage prioritization and selective human review, though they remain unsuitable for autonomous deployment.
📝 Abstract
Online patient inquiries are often informal, incomplete, and written before professional assessment, yet they must still be routed to an appropriate level of clinical follow-up. We study this as a four-class actionable triage task -- self-care, schedule-visit, urgent-clinician-review, or emergency-referral, and ask whether prompted large language models (LLMs) can support such routing under low-resource labeling conditions. Using the public HealthCareMagic-100K corpus, we construct a 300-example human calibrated gold evaluation set, a 700-example auto-labeled silver training set, and a 40-example few-shot pool. We compare Term Frequency-Inverse Document Frequency (TF-IDF) and Bidirectional Encoder Representations from Transformers for Biomedical Text Mining (BioBERT) baselines train on silver labels against six prompted LLMs under 0-shot, 4-shot, and 12-shot conditions respectively. Accordingly, we evaluate with macro-$F_1$ alongside safety-aware metrics, including emergency-recall, under-triage rate, and severe under-triage rate. The strongest LLM (Claude Haiku 4.5, 12-shot) reaches macro-$F_1$ 0.475, exceeding the best supervised baseline (BioBERT, 0.378) on point estimate, with overlapping confidence intervals. Few-shot prompting and two-model agreement help in label-dependent ways: self-care agreement is reliable, urgent-clinician-review is not. We conclude that LLMs can support triage prioritization and selective human review, but not autonomous deployment.
Problem

Research questions and friction points this paper is trying to address.

few-shot learning
large language models
triage categorization
online patient inquiries
low-resource labeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

few-shot learning
large language models
clinical triage
safety-aware evaluation
low-resource NLP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Liqi Zhou
J
Jiafu Li