CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that black-box visual grounding reasoning obscures intermediate spatial errors, hindering their diagnosis and correction. To this end, we propose a "construct-edit" framework that decouples explicit state construction from editing. Methodologically, we introduce a region evolution reinforcement mechanism and a bidirectional co-location reconstruction technique, employing a bidirectional denoising refiner to rectify coordinate trajectories. By integrating multimodal large language models with geometry- and behavior-level objective functions, our approach achieves interpretable and editable spatial state optimization. Experimental results demonstrate that a 9B-parameter model attains accuracy comparable to 241B-parameter counterparts. Furthermore, a single refinement step improves bounding box overlap by over 27%, substantially enhancing error-correction capabilities.
📝 Abstract
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.
Problem

Research questions and friction points this paper is trying to address.

Visual Grounding
Intermediate Localization Errors
Spatial Reasoning
Boundary Estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Grounding
Construct-to-Edit Framework
Region-Evolution Reinforcement
Bidirectional Denoising Refiner
State Refinement
💼 Related Jobs
No related jobs found.
D
Dongwei Sun
School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University, Xi’an 710049, China
Yujie Zhang
Yujie Zhang
Shanghai Jiao tong University
3D Quality AssessmentGeometry Processing3D Reconstruction
B
Bowen Yao
School of Computer Science and Technology, Faculty of Electronic and Information Engineering, Xi’an Jiaotong University, Xi’an 710049, China
Pei Liu
Pei Liu
The Hong Kong University of Science and Technoly
End-to-end Autonomous DrivingLarge Language Models
J
Jing Yao
State Key Laboratory of Remote Sensing and Digital Earth, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China
X
Xiangyong Cao
School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University, Xi’an 710049, China