Score
Designs, builds, and evaluates systems that use large language models to rank or prioritize candidate handlers (processing modules, remediation actions, or responders) for triage by estimating their relevance, risk, or likelihood of being problematic. Work includes prompt and interface design, feature extraction from inputs and handler descriptions, ranking or scoring models and heuristics, and evaluation/aggregation methods and human-in-the-loop workflows to present ranked handler lists to analysts at scale.
Software issue triage faces challenges of low operational efficiency and a persistent academic–industrial gap in complex system maintenance. To address this, we conduct a systematic literature review (SLR) spanning 234 English and Chinese publications from 2004 to 2023. This is the first SLR to jointly analyze academic research and industrial practice, revealing three critical bottlenecks: misaligned objectives between academia and industry, lack of standardized evaluation criteria, and practical deployment barriers. We propose a unified triage evaluation framework structured along four dimensions—data, tasks, metrics, and benchmarks—and systematically catalog open-source datasets and empirical methodologies to enable reproducible performance validation. All reviewed literature and supporting resources are publicly released. Our work establishes a foundational theory, provides actionable evaluation tools, and outlines collaborative pathways to bridge the gap between laboratory research and industrial adoption of triage technologies.
This study addresses the challenge of accurately triaging informal and often incomplete online patient queries into four actionable categories—self-care, scheduled appointment, urgent review, and emergency referral—under conditions of scarce labeled data. It presents the first real-world evaluation of few-shot large language models (LLMs) for clinical triage, combining human-calibrated and automatically annotated data to assess six LLMs alongside TF-IDF and BioBERT baselines across 0-shot to 12-shot settings. Results show that Claude Haiku 4.5 with 12 examples achieves a macro-F1 score of 0.475, significantly outperforming the best supervised baseline (BioBERT, 0.378). The work further introduces safety-aware metrics and dual-model consistency analysis, demonstrating that LLMs can support triage prioritization and selective human review, though they remain unsuitable for autonomous deployment.
This study empirically evaluates large language models (LLMs) against industry-standard technical hiring assessments for algorithm and software engineering roles. Method: We administered realistic, industrial-grade programming, system design, and reasoning questions—commonly used by leading technology firms—to state-of-the-art LLMs (e.g., GPT-4, Claude 3, Gemini) and conducted multi-stage comparative analysis against official corporate reference solutions, assessing correctness, completeness, engineering soundness, and consistency. Contribution/Results: Our analysis reveals systematic structural gaps between LLM outputs and industrial expectations: no tested model met enterprise hiring thresholds. Critical deficiencies were observed in boundary-case handling, explicit modeling of resource constraints (e.g., time/space complexity, scalability), and maintainability-aware design. These findings challenge the prevailing assumption that LLMs can directly substitute for entry-level engineers. Moreover, this work introduces the first benchmark framework specifically tailored to industrial recruitment scenarios, providing empirically grounded insights for AI capability evaluation in real-world engineering hiring.
This work addresses the limitation of current large language models (LLMs) in high-stakes clinical settings—such as emergency triage—where decisions entail asymmetric costs, like those of missed diagnoses versus over-triage. The study introduces a novel framework that embeds LLMs within a probabilistic decision-theoretic paradigm by incorporating configurable utility functions. This enables the model to dynamically generate triage recommendations aligned with explicit risk preferences, even when underlying diagnostic predictions remain unchanged. The approach integrates structured clinical case evaluation, utility-sensitive decision analysis, and probability calibration of model outputs. Experimental results demonstrate that the model’s recommendations are not solely driven by predictive accuracy but are effectively steered by explicit utility objectives, offering an interpretable, controllable, and goal-directed deployment paradigm for high-risk AI systems in clinical decision support.
This study systematically evaluates the robustness and fairness of large language models (LLMs) in emergency triage, focusing on distributional shift, missing data handling, and intersectional bias across gender and race. We propose a multi-strategy LLM triage evaluation framework integrating continual pretraining, in-context learning, and hybrid machine learning, augmented with counterfactual reasoning and robustness diagnostics. Our work is the first to empirically uncover significant intersectional bias in clinical triage: LLMs exhibit markedly reduced recommendation consistency for Black women—revealing implicit demographic preferences that may compromise real-world decision-making. Experiments demonstrate that LLMs outperform traditional models under data scarcity and distributional shift; however, their fairness deficiencies remain pronounced. This research provides critical empirical evidence and methodological foundations for robust deployment and bias mitigation of AI in healthcare.
This study addresses the limitations of conventional safety metrics in evaluating clinical triage large language models (LLMs) deployed in low-income settings, where apparent compliance may mask systemic unreliability. Leveraging real-world data from 19 primary care clinics in Nigeria, the authors construct IyawoBench v2.0—a benchmark comprising 200 synthetic cases—and introduce the first formal triage safety framework, which decomposes safety into three failure modes and includes 14 formal definitions and two theorems. They propose novel evaluation metrics, including the Escalation Bias Index and Expected Deployment Cost, revealing that all state-of-the-art models exhibit systemic failures. Notably, traditional sensitivity metrics obscure a 77-percentage-point undertriage gap in Llama 3.1 8B. The optimal model choice is shown to be highly dependent on deployment objectives—such as prioritizing emergency care, system sustainability, or balanced performance.
This study addresses the high variability in free-text triage notes within emergency departments, which contributes to inconsistent Emergency Severity Index (ESI) acuity assignments and compromises clinical accuracy and efficiency. To mitigate this issue, the authors propose an institution-tailored small language model (SLM) decision support system based on Qwen2.5-7B, fine-tuned through domain adaptation using expert-annotated data and a silver-standard pediatric triage dataset. The work systematically evaluates multiple prompting strategies and demonstrates that clinical-summary prompts yield the best performance. The fine-tuned model significantly reduces both ESI assignment inconsistency and clinically significant error rates, outperforming existing open-source SLMs as well as closed-source large models such as GPT-4o, while achieving a balanced integration of accuracy, stability, and patient privacy preservation.