🤖 AI Summary
This work addresses the prevalent confidence miscalibration and demographic unfairness of multimodal large language models (MLLMs) in few-shot medical image in-context learning. We propose CALIN, a fine-tuning-free inference-time calibration framework. CALIN introduces a novel two-level calibration matrix estimation: first constructing a population-level baseline calibration matrix from aggregate statistics, then dynamically adapting subgroup-specific calibration parameters along demographic dimensions (e.g., age, sex) for fine-grained confidence recalibration. To our knowledge, CALIN is the first to systematically uncover the coupled calibration–fairness bias of MLLMs across medical imaging benchmarks—including PAPILA, HAM10000, and MIMIC-CXR. Without degrading original task performance, CALIN significantly improves expected calibration error (ECE) and fairness metrics such as Equalized Odds difference across subgroups, achieving near-zero trade-off between fairness and utility.
📝 Abstract
Multimodal large language models (MLLMs) have enormous potential to perform few-shot in-context learning in the context of medical image analysis. However, safe deployment of these models into real-world clinical practice requires an in-depth analysis of the accuracies of their predictions, and their associated calibration errors, particularly across different demographic subgroups. In this work, we present the first investigation into the calibration biases and demographic unfairness of MLLMs' predictions and confidence scores in few-shot in-context learning for medical image classification. We introduce CALIN, an inference-time calibration method designed to mitigate the associated biases. Specifically, CALIN estimates the amount of calibration needed, represented by calibration matrices, using a bi-level procedure: progressing from the population level to the subgroup level prior to inference. It then applies this estimation to calibrate the predicted confidence scores during inference. Experimental results on three medical imaging datasets: PAPILA for fundus image classification, HAM10000 for skin cancer classification, and MIMIC-CXR for chest X-ray classification demonstrate CALIN's effectiveness at ensuring fair confidence calibration in its prediction, while improving its overall prediction accuracies and exhibiting minimum fairness-utility trade-off. Our codebase can be found at https://github.com/xingbpshen/medical-calibration-fairness-mllm.