🤖 AI Summary
This work addresses the limitation of current large language models (LLMs) in high-stakes clinical settings—such as emergency triage—where decisions entail asymmetric costs, like those of missed diagnoses versus over-triage. The study introduces a novel framework that embeds LLMs within a probabilistic decision-theoretic paradigm by incorporating configurable utility functions. This enables the model to dynamically generate triage recommendations aligned with explicit risk preferences, even when underlying diagnostic predictions remain unchanged. The approach integrates structured clinical case evaluation, utility-sensitive decision analysis, and probability calibration of model outputs. Experimental results demonstrate that the model’s recommendations are not solely driven by predictive accuracy but are effectively steered by explicit utility objectives, offering an interpretable, controllable, and goal-directed deployment paradigm for high-risk AI systems in clinical decision support.
📝 Abstract
High-stakes decisions under uncertainty, such as medical emergency triage, require more than accurate predictions. They depend on estimating the likelihood of alternative outcomes while explicitly weighing the consequences of different actions, principles that have long formed the foundation of medical diagnosis and decision making. Yet language models are increasingly used for high-stakes clinical recommendations without explicit specification of the utilities governing these decisions. Here we show that emergency triage with language models can be understood within a probabilistic decision framework, providing a case study of a broader decision-analytic paradigm for steering, evaluating, and deploying language models in high-stakes settings. Using clinical vignettes from a structured evaluation of a consumer triage system, we analyze recommendations for treatment under alternative utility functions that specify the relative costs of missed emergencies and unnecessary escalation. We find that capable language models adjust recommendations in response to stated utilities, revealing that the same underlying predictions can support markedly different decision policies. These findings show that effective deployment depends not only on improving predictions but also on making decision objectives explicit. More broadly, they suggest that language models for high-stakes applications should be understood and evaluated as probabilistic decision systems whose recommendations depend jointly on predictive performance and explicit utilities.