🤖 AI Summary
This study addresses the miscalibration between confidence and accuracy in medical multimodal large language models caused by text generation. To this end, it proposes the OmniMed-Jev framework, which introduces a decision-native interface that departs from the conventional next-token prediction paradigm. Instead, it uniformly formulates heterogeneous medical tasks as probability distributions over candidate sets, incorporating a System-1 thinking mechanism to yield trustworthy decision outputs. Experimental results demonstrate that, compared with generative baselines, the proposed method maintains comparable point-prediction performance while reducing calibration error by one order of magnitude and reliability error by two orders of magnitude. These findings indicate that OmniMed-Jev fundamentally achieves effective calibration of model confidence for reliable medical decision-making.
📝 Abstract
Medical models are judged not only on correctness, but on whether reported confidence matches actual accuracy. Generalist multimodal medical models have expanded what a single model can perceive, yet they still express bounded decisions such as diagnoses, findings or cell counts as generated text, so the reported probability reflects the next token rather than the decision itself. Motivated by decision-native interfaces such as Jev, we introduce OmniMed-Jev, which represents each medical decision as a Choice, Noul or Score decision over a runtime-supplied candidate set and returns a full distribution over that set: mutually exclusive classes, binary presence of a finding, or a bounded ordered value. The design is omni in three respects: it accepts diverse imaging modalities, covers different prediction tasks, and expresses them through one candidate-conditioned probability model, so heterogeneous outputs become comparable probabilities rather than task-specific strings. In an interface-controlled comparison against a generative baseline trained on the same backbone, data and schedule, OmniMed-Jev's reported probabilities track observed correctness far more closely, reducing calibration error by up to an order of magnitude and reliability error by up to two, while point-prediction performance remains comparable; counting is the one family where the generative baseline stays ahead. Making the decision distribution the model's output is not a format change but what turns reported numbers into probabilities that mean what they say. These results support explicit decision modeling as a way to make reported confidence meaningful within the evaluated tasks, and they are not evidence of clinical readiness: the comparison cannot separate the interface from associated training differences, which we state alongside the results. Code is available at github.com/lytang63/OmniMed-Jev.