π€ AI Summary
This study addresses the limitations of cascade architectures in generalized referring expression segmentation, where static interfaces and single-pass predictions hinder adaptive perception and geometric refinement. To overcome these issues, this work reformulates the task as a closed-loop visuo-motor process. Treating editable contours as states, it introduces an evolutionary perception semantic scheduling mechanism for multimodal feature routing. Furthermore, a dustbin-augmented entropy credit transport GRPO algorithm is designed, integrating bidirectional boundary sampling with instance-level reward modeling to jointly optimize discrete localization and continuous actions via reinforcement learning. Experiments demonstrate that the proposed method yields significant gIoU improvements across all gRefCOCO subsets and achieves state-of-the-art mIoU performance on all RefCOCO-series benchmarks.
π Abstract
Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation. This requirement challenges existing cascaded vision-language architectures, which typically rely on static feature interfaces and single-pass mask prediction, limiting adaptive perception and geometric correction. We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks. Evolution-Aware Semantic Scheduling (EASS) couples contour-guided bidirectional boundary sampling with state-conditioned routing of multilevel multimodal features, adapting perception to each contour state. Following supervised initialization, Dustbin-Augmented Entropic Credit Transport GRPO (DECT-GRPO) jointly optimizes discrete grounding and continuous contour actions with instance-level credits. Its rollout rewards and credits are derived from soft prediction-target correspondences that account for false positives and missed targets. ContourVLA improves gIoU over the strongest evaluated baselines by 8.7, 2.8, and 2.7 points on gRefCOCO val, testA, and testB, respectively, and achieves the highest mIoU across all eight RefCOCO, RefCOCO+, and RefCOCOg splits.