Score
Designs and implements models and pipelines that extract or generate structured semantic attributes and textual descriptions of visual content (e.g., category, color, shape, pose, visible parts, and context), including zero-shot descriptions produced by vision–language models. Builds and evaluates automated verification and alignment components—attribute classifiers, verification metrics, and filtering modules—that measure and enforce correctness and alignment between prompts, generated outputs, and extracted attributes.
Large language models (LLMs) inherently lack native visual understanding, limiting their applicability in multimodal scenarios. This work presents a systematic survey of vision-language models (VLMs), introducing— for the first time—a three-tier taxonomy grounded in modal input/output capabilities: comprehension-only, generation-only, and full-modality VLMs. We unify analysis across architectural design, training data composition, robustness properties, and benchmark performance (e.g., VQAv2, COCO Caption). Through comprehensive literature review, architectural decomposition, and cross-benchmark evaluation, we analyze over 100 state-of-the-art works to construct a technology evolution map. Our key contributions are: (1) the first scalable, capability-aware classification and evaluation framework for VLMs; (2) a precise delineation of current performance boundaries; and (3) identification of three critical future directions—embodied intelligence, robust multimodal reasoning, and efficient scaling—establishing an authoritative reference for the VLM research community.
This paper addresses the problem of evaluating the quality of textual descriptors—such as class names or descriptive phrases—in vision tasks. Existing methods rely excessively on classification accuracy and fail to characterize the intrinsic representational capacity of descriptors. To overcome this limitation, we propose a dual-dimensional evaluation framework: (1) a representation dimension, quantifying descriptor quality via two novel alignment metrics—global alignment and CLIP similarity—in the vision-language embedding space; and (2) a semantic compatibility dimension, measuring alignment with pretraining corpora of foundation models. We systematically benchmark mainstream descriptor generation strategies—including zero-shot LLM generation and iterative optimization—across VLMs such as CLIP. Experimental results demonstrate that our metrics effectively discriminate descriptor quality, uncover interactions between generation strategies and model architectures, and provide both theoretical foundations and practical tools for interpretable, scalable visual descriptor design.
Vision-language models (VLMs) suffer from low accuracy and poor generalization in business document chart understanding due to inherent visual recognition limitations. Method: We propose a text-only paradigm for chart structure understanding—bypassing image-based analysis entirely and instead parsing structured metadata (e.g., shapes, connectors) directly from editable source files (XLSX/PPTX/DOCX) at the Office Open XML (OOXML) level, then feeding this structured input to large language models (LLMs) for relational reasoning and question answering. Contribution/Results: By eliminating VLMs’ visual bottlenecks and leveraging fine-grained XML parsing with structure-aware prompt engineering, our approach achieves high-precision semantic parsing. On system design document QA tasks, it significantly outperforms VLM baselines. Robust cross-format generalization is validated across PPTX, DOCX, and XLSX, demonstrating strong adaptability to real-world business scenarios. This work establishes a new, interpretable, cost-effective, and high-accuracy pathway for document intelligence.
This work addresses the challenge of fine-grained, structured annotation of human body language—including pose and emotion—in video. We propose an end-to-end pipeline leveraging dual vision-language models (VLMs): Qwen2.5-VL-7B and Llama-4-Scout-17B. Our method integrates visual tokenization, multimodal Transformer attention, and instruction tuning to achieve frame-level person detection (pixel-accurate bounding boxes), prompt-conditioned emotion recognition, and cross-frame ID-consistent modeling, augmented by a schema-driven output validation module ensuring structural compliance. Methodologically, we are the first to systematically disentangle critical boundaries—syntactic validity versus semantic correctness, structural validation versus geometric precision, and local frame-level ID assignment versus cross-frame tracking—explicitly guided by VLM architectural properties. Experiments demonstrate reproducibility, interface robustness, and evaluation reliability, establishing a novel paradigm for controllable VLM deployment in embodied perception tasks.
Existing vision-language prompt learning methods merely concatenate learnable prompts with class names, neglecting the rich semantic context embedded in class names. Method: We propose TreePrompt, which (1) leverages large language models to generate a hierarchical “concept–attribute–description” knowledge tree to explicitly model fine-grained visual attributes; (2) introduces a vision-conditioned text pooling module to achieve instance-level image–text prompt alignment; and (3) elevates prompt learning into interpretable, structured domain-expert modeling. Our approach integrates tree-structured knowledge distillation, hierarchical vision–language co-learning, and structured knowledge graph embedding. Contribution/Results: TreePrompt achieves state-of-the-art performance across 11 benchmarks, delivering significant improvements in zero-shot base-to-novel class generalization, cross-dataset transfer, and few-shot classification.
This work addresses the challenge of simultaneously achieving structural precision, semantic interpretability, and identity controllability in existing 3D/4D scene representations. We propose “Scene Language”—a unified 3D/4D scene representation framework that integrates executable program structures, natural-language semantic tokens, and visual identity embeddings. To our knowledge, this is the first method enabling zero-shot cross-modal reasoning: without fine-tuning, it directly synthesizes structured scene programs from pretrained language models and vision encoders, while explicitly modeling hierarchical relationships to support fine-grained editing. The representation is renderer-agnostic, interfacing seamlessly with traditional, neural, and hybrid renderers to produce high-fidelity images. Experiments demonstrate significant improvements over baselines—including scene graphs—on complex scene generation tasks, achieving breakthroughs in fidelity, controllability, and editability.
When BPMN process diagram source files (e.g., XML) are unavailable, recovering structured semantic information directly from raster images remains challenging. Method: We propose an end-to-end vision-language joint approach that tightly integrates multimodal large models (VLMs) with optical character recognition (OCR) via prompt engineering—enabling unified modeling of graphical symbol recognition, text localization, and semantic alignment without manual annotations or textual priors. Contribution/Results: Ablation studies and statistical analysis across multiple VLM benchmarks demonstrate that OCR enhancement significantly improves node-type identification and control-flow relation extraction accuracy (average +12.7%). The method exhibits strong robustness against image degradation—including blurriness, scaling artifacts, and low resolution. This work establishes a practical, deployable paradigm for structured image parsing in reverse engineering and legacy system digitization.
Generative models face severe challenges in transparency and copyright traceability, primarily due to the difficulty of establishing fine-grained provenance links between generated content and original training data. Method: This paper proposes an ontology-aligned knowledge graph construction method leveraging multimodal large language models (MLLMs). It first parses semantic content from generated images and extracts structured subject–predicate–object triples; then performs cross-modal and cross-source ontology alignment to unify heterogeneous knowledge representations; finally enables interpretable, sample-level溯源 from outputs back to training instances. Contribution/Results: The method is validated on both local and large-scale models, significantly improving copyright attribution accuracy and dataset transparency. It provides a scalable, technically grounded foundation for responsible governance of generative AI and human-AI collaboration, advancing traceability beyond black-box generation.
To address the challenges of high-resolution image synthesis and multimodal semantic understanding, this paper introduces VLM-RF, a vision-enhanced large language model. Methodologically, it pioneers a noise-aware learning algorithm and integrates a linear-path rectified flow (RF) mechanism with a cross-modal bidirectional tokenization strategy, enabling unified spatiotemporal feature embedding and hybrid sequence modeling across text, images, and video. The contributions are threefold: (1) substantial improvement in generation quality—image resolution and perceptual sharpness increase by 25%; (2) 20% reduction in computational overhead; and (3) consistent superiority over state-of-the-art diffusion models in both synthesis fidelity and cross-modal alignment. By unifying generative modeling and language understanding within a single scalable architecture, VLM-RF establishes a novel paradigm for efficient, high-fidelity multimodal generation.
This work addresses two critical challenges in generative zero-shot learning: the class-instance gap caused by intra-class variation and the domain gap arising from misalignment between semantic and visual feature distributions. To tackle these issues, the authors propose a unified modeling and alignment framework. It incorporates an Attribute Distribution Modeling (ADM) module to learn transferable class-level attribute distributions and sample instance-level attributes, along with a Visual-Guided Alignment (VGA) module that leverages visual information to refine semantic representations and explicitly align semantic and visual spaces. The proposed method achieves significant performance gains, improving state-of-the-art results by 4.7% on AWA2 and 6.1% on SUN benchmarks. Furthermore, it functions as a plug-and-play component that can effectively enhance other generative ZSL models.
This work addresses the challenges faced by general-purpose vision-language models (VLMs) in e-commerce settings, where dense attribute spaces, multi-image inputs, and noisy data hinder simultaneous optimization of domain-specific adaptation and general multimodal capabilities. To tackle this, the authors propose an e-commerce-oriented VLM adaptation strategy that incorporates multi-image fusion and structured attribute modeling through targeted fine-tuning. They further introduce a comprehensive evaluation framework encompassing deep product understanding, instruction following, and dynamic attribute extraction. Experimental results demonstrate that the proposed approach significantly enhances model performance on e-commerce tasks while effectively preserving its general multimodal generalization ability.