Score
Designs, implements, and evaluates models and algorithms that link semantic symbols or natural language expressions to perceptual or structured referents — for example image regions, segmentation masks, spatial coordinates, or entries in an intermediate semantic layer — so that linguistic queries and labels can be resolved to concrete parts of an input. Builds segmentation-aware region encoders, region-based and dense grounding networks, and mediation/abstraction layers that map between ontologies and spatial or structured representations for downstream resolution, planning, retrieval, or reasoning.
This survey addresses the lack of systematic reviews and standardized definitions in visual grounding. It comprehensively surveys the past decade’s research, unifying foundational task definitions and diverse settings—including grounding pretraining, MLLM-based grounding, and generalized/ultra-high-resolution grounding—and establishes a standardized evaluation framework. Methodologically, it synthesizes advances in deep learning, cross-modal alignment, contrastive learning, Transformers, and multimodal large language models (MLLMs), analyzing evolutionary trends and core challenges. Key contributions include: (1) constructing the most comprehensive knowledge system for visual grounding to date; (2) open-sourcing a high-quality resource repository (Awesome-Visual-Grounding); (3) proposing emerging directions such as grounded MLLMs and giga-pixel-level localization; and (4) releasing a structured knowledge graph tailored for both beginners and experts—thereby significantly advancing field standardization and sustainable development.
Existing referring expression segmentation methods employ a uniform refinement strategy that struggles to balance diverse semantic and spatial reasoning demands, often resulting in fragmented outputs, ambiguous boundaries, or misidentified targets—particularly when the pretrained backbone is frozen. To address this, this work proposes SERA, a novel architecture featuring a two-stage, lightweight, expression-aware expert refinement mechanism. SERA integrates an expression-conditioned adapter (SERA-Adapter) into the backbone to enhance spatial consistency and boundary precision, and introduces a geometry-preserving expert transformation (SERA-Fusion) applied to visual features prior to multimodal fusion. By combining mixture-of-experts with a lightweight routing scheme, SERA achieves substantial improvements in spatial localization and boundary delineation while fine-tuning fewer than 1% of parameters—limited to normalization layers and bias terms—consistently outperforming strong baselines on standard benchmarks, especially in high-precision spatial reference tasks.
Aerial image referring expression segmentation faces challenges including large resolution variations, inconsistent color characteristics, small and densely packed objects, and partial occlusions. To address these, we propose Aerial-D, the first large-scale, cross-era aerial image referring expression segmentation dataset, comprising over 1.5 million diverse language expressions generated via a hybrid pipeline integrating rule-based systems and large language models. We further introduce an imaging degradation simulation filter to adapt modern models to historical imagery (e.g., grayscale, sepia-toned, or grainy). Our method extends the RSRefSeg architecture to enable unified segmentation modeling for both contemporary and historical aerial images. Experiments demonstrate state-of-the-art performance on modern benchmarks and exceptional robustness across multiple synthetic degradation conditions. Aerial-D is the first dataset and framework enabling precise, text-driven, multi-temporal aerial image segmentation.
Existing referring expression segmentation and grounding methods rely on single-sentence descriptions, failing to adequately capture visually rich object details and thus suffering from misidentification of similar objects. To address this, we propose a vision-enhanced latent expression generation framework. First, a subject allocation and visual concept injection module synthesizes multiple diverse, attribute-rich latent textual expressions from a single input sentence. Second, a positive-margin contrastive learning mechanism is introduced to model fine-grained distinctions while preserving semantic consistency. Finally, leveraging a shared-subject–distinct-attribute disentangled latent space, we jointly optimize cross-modal alignment and text–latent expression co-refinement. Our method achieves state-of-the-art performance on multiple referring expression segmentation and comprehension benchmarks, and significantly outperforms prior work on the generalized referring expression segmentation (GRES) task.
This study aims to identify the “semantic core layers”—those encoding the richest semantic information—in large language models (LLMs) and vision Transformers (ViTs), and to quantify the information content and directional asymmetry of cross-modal representations. Method: We propose a quantitative analytical framework grounded in the information bottleneck principle and inter-layer mutual information, integrated with translation-aligned modeling, caption–image prediction, and cross-modal similarity probing. Contribution/Results: We systematically identify semantic-critical layers in LLMs (DeepSeek-V3, Llama3.1-8B) and ViTs for the first time. Key findings include: (i) semantic information exhibits long-range token dependencies and causal asymmetry across layers; (ii) semantic layers in LLMs generalize to predict ViT image representations; and (iii) cross-modal semantic information flow displays strong, model-dependent unidirectional dominance—e.g., from text to image or vice versa—rather than symmetry. These results provide theoretical foundations and a novel interpretability pathway for multimodal representation learning.
This work investigates the capacity of large language models (LLMs) to perform spatial semantic understanding and cross-modal reasoning solely from symbolic graphical programs—such as curve parameters, stroke sequences, and local curvature—without visual encoders. To this end, we introduce the first benchmark for symbolic graphical program-based visual understanding, comprising three tasks: program generation, semantic question answering, and zero-visual-input cross-modal reasoning. We propose Symbolic Instruction Tuning (SIT), a novel fine-tuning paradigm that explicitly enhances LLMs’ spatial reasoning capabilities using synthetic symbolic graphical instruction data. Experimental results demonstrate that strong reasoning-oriented LLMs achieve superior performance; SIT substantially improves accuracy on symbolic graphical understanding tasks and—unexpectedly—generalizes to multiple general-purpose reasoning benchmarks (e.g., GSM8K, MMLU), indicating that symbolic spatial representations can strengthen foundational reasoning abilities.
This work addresses the challenge faced by visually impaired users in accessing node-link diagrams commonly distributed as bitmap images, a task for which existing assistive technologies are ill-suited due to their reliance on structured data rather than visual input. The paper presents the first lightweight deep learning approach for semantic segmentation of such diagram images, training a compact model on a large-scale synthetic dataset to achieve pixel-level parsing. The proposed method attains over 93% pixel accuracy on synthetic data and demonstrates strong performance both quantitatively and qualitatively. By enabling precise extraction of diagram semantics directly from rasterized images, this approach establishes a viable foundation for non-visual interaction and effectively bridges a critical gap in accessibility technology for bitmap-based graphical content.
Current multimodal large language models (MLLMs) rely on coarse-grained bounding boxes for region-focused visual reasoning, which are prone to background interference and misaligned with the underlying visual token structure. This work proposes SegAnswer, the first approach to integrate pixel-level segmentation masks into the MLLM inference pipeline, replacing bounding boxes with fine-grained masks as the visual focus units. SegAnswer leverages a segmentation model to generate masks, crops the input image accordingly, and fuses positional embeddings to achieve precise region isolation and natural alignment with visual tokens. Experiments demonstrate consistent improvements across tasks requiring high-resolution perception, general visual understanding, and hallucination suppression, while also exhibiting robust pixel-level localization capabilities.
This work addresses the persistent limitations of large language models (LLMs) in understanding real-world concepts, particularly their insufficient grasp of core conceptual attributes. To this end, the paper introduces the first controllable evaluation benchmark focused on geospatial concepts—such as direction, distance, and topology—employing synthetic data–driven question-answering tasks to systematically assess LLMs across three dimensions: abstractness, compositionality, and grounding. The study reveals significant deficiencies in current models’ ability to acquire and compose structured conceptual knowledge, while also elucidating how model scale and architecture influence conceptual understanding. These findings offer critical insights for the future design of more cognitively capable language models.
This study addresses the difficulty of Vision Transformers in effectively grounding abstract concepts without direct visual referents. To this end, it proposes a “metaphorical anchoring” mechanism that leverages concrete, interpretable concepts to bridge visual signals and abstract semantics. Methodologically, Transcoders are applied to parse intermediate features from CLIP and DINO encoders, combined with circuit tracing over an icon dataset. The analysis reveals structured metaphorical circuits wherein perceptual primitives dominate early layers and object-level anchors emerge prior to abstract targets. Furthermore, causal interventions confirm the functional involvement of mediating concepts in abstract concept grounding. Collectively, these findings establish a novel paradigm for understanding abstract reasoning within vision-language models.
This work addresses the challenge that current vision-language models struggle to emulate humans’ ability to develop shared referential expressions through repeated interaction—a phenomenon known as lexical coordination. To bridge this gap, the authors propose a dynamic semantic framework that explicitly maintains three sets of reference-object binding states and integrates a lightweight perceptual alignment module leveraging SIFT homography, Universal Quality Index (UQI), and image augmentation techniques. This architecture externalizes the lexical coordination process into an inspectable symbolic layer, yielding a transparent and auditable mechanism for referential understanding amenable to fine-grained ablation studies. Evaluated on the Stanford Referring Expression Repetition Game corpus, the model achieves 83.56% accuracy in identifying the target within the top-5 candidates using only a single utterance, and demonstrates robust performance even under conservative evaluation conditions that exclude salient distractors.