Disentangling Prompt Dependence to Evaluate Segmentation Reliability in Gynecological MRI

📅 2026-03-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliability issues of promptable segmentation models in gynecological MRI arising from user-dependent prompt variations. The authors propose an interpretable framework that, for the first time, disentangles prompt dependency into two distinct components: prompt ambiguity (inter-user variability) and local sensitivity (imprecision in user interaction). Built upon the Segment Anything Model, the framework employs quantitative metrics to analyze the relationship between prompt variability and segmentation performance for uterine and bladder delineation. Evaluated on two female pelvic MRI datasets, the proposed metrics exhibit strong negative correlations with segmentation accuracy while demonstrating low mutual correlation, thereby effectively revealing distinct prompt-related failure modes. This approach provides a principled basis for evaluating model robustness and supports safer clinical deployment of interactive segmentation systems.

Technology Category

Application Category

📝 Abstract
Promptable segmentation models (e.g., the Segment Anything Models) enable generalizable, zero-shot segmentation across diverse domains. Although predictions are deterministic for a fixed image-prompt pair, the robustness of these models to variations in user prompts, referred to as prompt dependence, remains underexplored. In safety-critical workflows with substantial inter-user variability, interpretable and informative frameworks are needed to evaluate prompt dependence. In this work, we assess the reliability of promptable segmentation by analyzing and measuring its sensitivity to prompt variability. We introduce the first formulation of prompt dependence that explicitly disentangles prompt ambiguity (inter-user variability) from local sensitivity (interaction imprecision), offering an interpretable view of segmentation robustness. Experiments on two female pelvic MRI datasets for uterus and bladder segmentation reveal a strong negative correlation between both metrics and segmentation performance, highlighting the value of our framework for assessing robustness. The two metrics have low mutual correlation, supporting the disentangled design of our formulation, and provide meaningful indicators of prompt-related failure modes.
Problem

Research questions and friction points this paper is trying to address.

prompt dependence
segmentation reliability
promptable segmentation
gynecological MRI
user variability
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt dependence
disentanglement
segmentation reliability
promptable segmentation
medical image analysis
🔎 Similar Papers
No similar papers found.
Elodie Germani
Elodie Germani
Universitätsklinikum Bonn
NeuroimagingfMRImachine learningreproducibilitystatistics
K
Krystel Nyangoh-Timoh
1 Laboratoire Traitement du Signal et de l’Image (LTSI, INSERM UMR 1099), Université de Rennes, Rennes, France; 2 Centre Hospitalier Universitaire de Rennes (CHU Rennes), Rennes, France
Pierre Jannin
Pierre Jannin
MediCIS, LTSI, Inserm, Université de Rennes
Surgical data scienceComputer Assisted Surgery
J
John S H Baxter
Laboratoire Traitement du Signal et de l’Image (LTSI, INSERM UMR 1099), Université de Rennes, Rennes, France