InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of existing referring segmentation evaluations to decouple reasoning logic, which impedes isolating target selection errors from mask generation errors. We construct an instance-centric diagnostic benchmark featuring a target-oriented expression set, a compact taxonomy of logical categories, and identity-aware metrics. By integrating minimal-difference pairing with key-construct suppression to mitigate shortcut learning, we establish a closed-loop "evaluate-diagnose-improve" framework. Evaluations across multiple models reveal that target selection constitutes the primary bottleneck. Furthermore, controlled intervention experiments demonstrate that matching supervision significantly enhances identity-aware performance, with the optimal model achieving an mIoU of 67.1%.
📝 Abstract
Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% All@0.7. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.
Problem

Research questions and friction points this paper is trying to address.

Referring Expression Segmentation
Referential Reasoning
Target Identity
Instance-level Evaluation
Error Decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Referring Expression Segmentation
Diagnostic Benchmark
Referential Reasoning
Identity-aware Metrics
Failure Decomposition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuchen Li
The University of Sydney
S
Shaoyang Zhou
The University of Sydney
Y
Yiran Wang
The University of Sydney
R
Ruiyi Deng
The University of Sydney
H
Haoyu Wang
The University of Sydney
Z
Ziru Wei
The ATLAS Institute
Z
Zhen Zhao
Shanghai Artificial Intelligence Laboratory
Luping Zhou
Luping Zhou
School of Electrical and Computer Engineering, University of Sydney
Medical ImagingComputer VisionMachine Learning