🤖 AI Summary
This study addresses the inability of existing referring segmentation evaluations to decouple reasoning logic, which impedes isolating target selection errors from mask generation errors. We construct an instance-centric diagnostic benchmark featuring a target-oriented expression set, a compact taxonomy of logical categories, and identity-aware metrics. By integrating minimal-difference pairing with key-construct suppression to mitigate shortcut learning, we establish a closed-loop "evaluate-diagnose-improve" framework. Evaluations across multiple models reveal that target selection constitutes the primary bottleneck. Furthermore, controlled intervention experiments demonstrate that matching supervision significantly enhances identity-aware performance, with the optimal model achieving an mIoU of 67.1%.
📝 Abstract
Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% All@0.7. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.