🤖 AI Summary
This paper systematically surveys visual grounding—the capability of vision-language models (VLMs) to precisely localize image regions corresponding to textual descriptions. Adopting a comprehensive literature review methodology, it analyzes prevailing techniques, application scenarios, and evaluation frameworks using standard benchmarks (e.g., RefCOCO, Flickr30K Entities) and metrics. The study establishes the first holistic research framework for visual grounding, explicitly characterizing its definition, methodologies, evaluation protocols, and open challenges—while uncovering its intrinsic connections to multimodal chain-of-thought reasoning and inferential capacity. Key limitations are identified in fine-grained localization accuracy, cross-dataset generalization, and model interpretability. To address these, the paper proposes three future directions: integrating explicit spatial modeling, reasoning-guided grounding, and unified evaluation standards. This work provides both theoretical foundations and practical guidance for advancing fine-grained cross-modal understanding.
📝 Abstract
Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in various domains, including referring expression comprehension, answering questions pertinent to fine-grained details in images or videos, caption visual context by explicitly referring to entities, as well as low and high-level control in simulated and real environments. In this survey paper, we review representative works across the key areas of research on modern general-purpose vision language models (VLMs). We first outline the importance of grounding in VLMs, then delineate the core components of the contemporary paradigm for developing grounded models, and examine their practical applications, including benchmarks and evaluation metrics for grounded multimodal generation. We also discuss the multifaceted interrelations among visual grounding, multimodal chain-of-thought, and reasoning in VLMs. Finally, we analyse the challenges inherent to visual grounding and suggest promising directions for future research.