ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of cascade architectures in generalized referring expression segmentation, where static interfaces and single-pass predictions hinder adaptive perception and geometric refinement. To overcome these issues, this work reformulates the task as a closed-loop visuo-motor process. Treating editable contours as states, it introduces an evolutionary perception semantic scheduling mechanism for multimodal feature routing. Furthermore, a dustbin-augmented entropy credit transport GRPO algorithm is designed, integrating bidirectional boundary sampling with instance-level reward modeling to jointly optimize discrete localization and continuous actions via reinforcement learning. Experiments demonstrate that the proposed method yields significant gIoU improvements across all gRefCOCO subsets and achieves state-of-the-art mIoU performance on all RefCOCO-series benchmarks.
πŸ“ Abstract
Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation. This requirement challenges existing cascaded vision-language architectures, which typically rely on static feature interfaces and single-pass mask prediction, limiting adaptive perception and geometric correction. We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks. Evolution-Aware Semantic Scheduling (EASS) couples contour-guided bidirectional boundary sampling with state-conditioned routing of multilevel multimodal features, adapting perception to each contour state. Following supervised initialization, Dustbin-Augmented Entropic Credit Transport GRPO (DECT-GRPO) jointly optimizes discrete grounding and continuous contour actions with instance-level credits. Its rollout rewards and credits are derived from soft prediction-target correspondences that account for false positives and missed targets. ContourVLA improves gIoU over the strongest evaluated baselines by 8.7, 2.8, and 2.7 points on gRefCOCO val, testA, and testB, respectively, and achieves the highest mIoU across all eight RefCOCO, RefCOCO+, and RefCOCOg splits.
Problem

Research questions and friction points this paper is trying to address.

Generalized Referring Expression Segmentation
Vision-Language Architecture
Adaptive Perception
Geometric Correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop Perception-Action Policy
Generalized Referring Expression Segmentation
Evolution-Aware Semantic Scheduling
DECT-GRPO
Editable Contours
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Ruicheng Zhang
Ruicheng Zhang
Sun Yat-sen University
ε€šζ¨‘ζ€ε€§ζ¨‘εž‹γ€ε…·θΊ«ζ™Ίθƒ½γ€εΌΊεŒ–ε­¦δΉ γ€εŒ»ε­¦ε›Ύεƒ
K
Kaiwen Shen
Sun Yat-sen University
Jiaqi Hou
Jiaqi Hou
Tsinghua University
S
Shuhan Yang
Tsinghua University
J
Junchao Huang
The Chinese University of Hong Kong, Shenzhen
K
Kewei Zhang
Peking University
J
Jun Zhou
Tsinghua University
Li Jiang
Li Jiang
The Chinese University of Hong Kong (Shenzhen)
Computer VisionDeep Learning
S
Shen Zhao
Sun Yat-sen University