semantic attribute extraction

Designs and implements models and pipelines that extract or generate structured semantic attributes and textual descriptions of visual content (e.g., category, color, shape, pose, visible parts, and context), including zero-shot descriptions produced by vision–language models. Builds and evaluates automated verification and alignment components—attribute classifiers, verification metrics, and filtering modules—that measure and enforce correctness and alignment between prompts, generated outputs, and extracted attributes.

semanticattributeextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Beyond Accuracy: Metrics that Uncover What Makes a `Good' Visual Descriptor

Jul 04, 2025
EL
Ethan Lin
🏛️ Cornell University | University of Texas at Austin

This paper addresses the problem of evaluating the quality of textual descriptors—such as class names or descriptive phrases—in vision tasks. Existing methods rely excessively on classification accuracy and fail to characterize the intrinsic representational capacity of descriptors. To overcome this limitation, we propose a dual-dimensional evaluation framework: (1) a representation dimension, quantifying descriptor quality via two novel alignment metrics—global alignment and CLIP similarity—in the vision-language embedding space; and (2) a semantic compatibility dimension, measuring alignment with pretraining corpora of foundation models. We systematically benchmark mainstream descriptor generation strategies—including zero-shot LLM generation and iterative optimization—across VLMs such as CLIP. Experimental results demonstrate that our metrics effectively discriminate descriptor quality, uncover interactions between generation strategies and model architectures, and provide both theoretical foundations and practical tools for interpretable, scalable visual descriptor design.

Analyzing visual descriptor quality beyond accuracy metricsEvaluating descriptor generation methods for vision-language modelsIntroducing alignment-based metrics to assess descriptor effectiveness

Vision-language models (VLMs) suffer from low accuracy and poor generalization in business document chart understanding due to inherent visual recognition limitations. Method: We propose a text-only paradigm for chart structure understanding—bypassing image-based analysis entirely and instead parsing structured metadata (e.g., shapes, connectors) directly from editable source files (XLSX/PPTX/DOCX) at the Office Open XML (OOXML) level, then feeding this structured input to large language models (LLMs) for relational reasoning and question answering. Contribution/Results: By eliminating VLMs’ visual bottlenecks and leveraging fine-grained XML parsing with structure-aware prompt engineering, our approach achieves high-precision semantic parsing. On system design document QA tasks, it significantly outperforms VLM baselines. Robust cross-format generalization is validated across PPTX, DOCX, and XLSX, demonstrating strong adaptability to real-world business scenarios. This work establishes a new, interpretable, cost-effective, and high-accuracy pathway for document intelligence.

Bypasses visual recognition limitations of Vision-Language ModelsEnhances diagram understanding with text-driven methodUtilizes source file metadata for accurate structure analysis

This work addresses the challenge of fine-grained, structured annotation of human body language—including pose and emotion—in video. We propose an end-to-end pipeline leveraging dual vision-language models (VLMs): Qwen2.5-VL-7B and Llama-4-Scout-17B. Our method integrates visual tokenization, multimodal Transformer attention, and instruction tuning to achieve frame-level person detection (pixel-accurate bounding boxes), prompt-conditioned emotion recognition, and cross-frame ID-consistent modeling, augmented by a schema-driven output validation module ensuring structural compliance. Methodologically, we are the first to systematically disentangle critical boundaries—syntactic validity versus semantic correctness, structural validation versus geometric precision, and local frame-level ID assignment versus cross-frame tracking—explicitly guided by VLM architectural properties. Experiments demonstrate reproducibility, interface robustness, and evaluation reliability, establishing a novel paradigm for controllable VLM deployment in embodied perception tasks.

Address semantic correctness and system constraints in video-to-artifact pipelinesDetect visible people and emotions from video frames using vision-language modelsGenerate structured bounding box outputs with prompt-conditioned attributes

Tree of Attributes Prompt Learning for Vision-Language Models

Oct 15, 2024
TD
Tong Ding
🏛️ Harvard University | Mass General Brigham | Microsoft

Existing vision-language prompt learning methods merely concatenate learnable prompts with class names, neglecting the rich semantic context embedded in class names. Method: We propose TreePrompt, which (1) leverages large language models to generate a hierarchical “concept–attribute–description” knowledge tree to explicitly model fine-grained visual attributes; (2) introduces a vision-conditioned text pooling module to achieve instance-level image–text prompt alignment; and (3) elevates prompt learning into interpretable, structured domain-expert modeling. Our approach integrates tree-structured knowledge distillation, hierarchical vision–language co-learning, and structured knowledge graph embedding. Contribution/Results: TreePrompt achieves state-of-the-art performance across 11 benchmarks, delivering significant improvements in zero-shot base-to-novel class generalization, cross-dataset transfer, and few-shot classification.

Addresses misalignment with vision-conditional text feature extractionEnhances vision-language models with structured attribute treesImproves text and vision prompt learning for attributes

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

Oct 22, 2024
YZ
Yunzhi Zhang
🏛️ Stanford University | UC Berkeley

This work addresses the challenge of simultaneously achieving structural precision, semantic interpretability, and identity controllability in existing 3D/4D scene representations. We propose “Scene Language”—a unified 3D/4D scene representation framework that integrates executable program structures, natural-language semantic tokens, and visual identity embeddings. To our knowledge, this is the first method enabling zero-shot cross-modal reasoning: without fine-tuning, it directly synthesizes structured scene programs from pretrained language models and vision encoders, while explicitly modeling hierarchical relationships to support fine-grained editing. The representation is renderer-agnostic, interfacing seamlessly with traditional, neural, and hybrid renderers to produce high-fidelity images. Experiments demonstrate significant improvements over baselines—including scene graphs—on complex scene generation tasks, achieving breakthroughs in fidelity, controllability, and editability.

Enabling high-quality 3D and 4D scene generation and editingInferring scene representation from text or image inputsRepresenting visual scenes with structure, semantics, and identity

Latest Papers

What's happening recently
View more

Structured Extraction from Business Process Diagrams Using Vision-Language Models

Nov 27, 2025
PD
Pritam Deka
🏛️ Queen’s University Belfast

When BPMN process diagram source files (e.g., XML) are unavailable, recovering structured semantic information directly from raster images remains challenging. Method: We propose an end-to-end vision-language joint approach that tightly integrates multimodal large models (VLMs) with optical character recognition (OCR) via prompt engineering—enabling unified modeling of graphical symbol recognition, text localization, and semantic alignment without manual annotations or textual priors. Contribution/Results: Ablation studies and statistical analysis across multiple VLM benchmarks demonstrate that OCR enhancement significantly improves node-type identification and control-flow relation extraction accuracy (average +12.7%). The method exhibits strong robustness against image degradation—including blurriness, scaling artifacts, and low resolution. This work establishes a practical, deployable paradigm for structured image parsing in reverse engineering and legacy system digitization.

Enrich extraction using OCR and evaluate accuracyExtract structured JSON from BPMN diagram imagesLeverage Vision-Language Models without source files

Training Data Attribution for Image Generation using Ontology-Aligned Knowledge Graphs

Dec 02, 2025
TA
Theodoros Aivalis
🏛️ National Centre for Scientific Research "Demokritos" | University of Glasgow

Generative models face severe challenges in transparency and copyright traceability, primarily due to the difficulty of establishing fine-grained provenance links between generated content and original training data. Method: This paper proposes an ontology-aligned knowledge graph construction method leveraging multimodal large language models (MLLMs). It first parses semantic content from generated images and extracts structured subject–predicate–object triples; then performs cross-modal and cross-source ontology alignment to unify heterogeneous knowledge representations; finally enables interpretable, sample-level溯源 from outputs back to training instances. Contribution/Results: The method is validated on both local and large-scale models, significantly improving copyright attribution accuracy and dataset transparency. It provides a scalable, technically grounded foundation for responsible governance of generative AI and human-AI collaboration, advancing traceability beyond black-box generation.

Attributing training data influence on generated imagesEnhancing transparency and copyright analysis in generative modelsExtracting structured knowledge from visual content using LLMs

To address the challenges of high-resolution image synthesis and multimodal semantic understanding, this paper introduces VLM-RF, a vision-enhanced large language model. Methodologically, it pioneers a noise-aware learning algorithm and integrates a linear-path rectified flow (RF) mechanism with a cross-modal bidirectional tokenization strategy, enabling unified spatiotemporal feature embedding and hybrid sequence modeling across text, images, and video. The contributions are threefold: (1) substantial improvement in generation quality—image resolution and perceptual sharpness increase by 25%; (2) 20% reduction in computational overhead; and (3) consistent superiority over state-of-the-art diffusion models in both synthesis fidelity and cross-modal alignment. By unifying generative modeling and language understanding within a single scalable architecture, VLM-RF establishes a novel paradigm for efficient, high-fidelity multimodal generation.

Enhance high-resolution image synthesis using vision-augmented LLMsImprove generative performance under noisy input conditionsInterpret multimodal data via unified text-image-video tokenization

This work addresses two critical challenges in generative zero-shot learning: the class-instance gap caused by intra-class variation and the domain gap arising from misalignment between semantic and visual feature distributions. To tackle these issues, the authors propose a unified modeling and alignment framework. It incorporates an Attribute Distribution Modeling (ADM) module to learn transferable class-level attribute distributions and sample instance-level attributes, along with a Visual-Guided Alignment (VGA) module that leverages visual information to refine semantic representations and explicitly align semantic and visual spaces. The proposed method achieves significant performance gains, improving state-of-the-art results by 4.7% on AWA2 and 6.1% on SUN benchmarks. Furthermore, it functions as a plug-and-play component that can effectively enhance other generative ZSL models.

attribute distributiondomain gapintra-class variability

This work addresses the challenges faced by general-purpose vision-language models (VLMs) in e-commerce settings, where dense attribute spaces, multi-image inputs, and noisy data hinder simultaneous optimization of domain-specific adaptation and general multimodal capabilities. To tackle this, the authors propose an e-commerce-oriented VLM adaptation strategy that incorporates multi-image fusion and structured attribute modeling through targeted fine-tuning. They further introduce a comprehensive evaluation framework encompassing deep product understanding, instruction following, and dynamic attribute extraction. Experimental results demonstrate that the proposed approach significantly enhances model performance on e-commerce tasks while effectively preserving its general multimodal generalization ability.

Attribute ExtractionE-commerceModel Adaptation

Hot Scholars

WL

Weihua Luo

Alibaba
natural language processingmachine learningartificial intelligence
LW

Longyue Wang

Alibaba International
Large Language ModelMachine TranslationNatural Language ProcessingLanguange Agent
SM

Shervin Malmasi

Harvard Medical School
Medical InformaticsNatural Language ProcessingComputational LinguisticsMachine Learning
KZ

Kaifu Zhang

Assistant Professor of Marketing, Carnegie Mellon University
Two-sided marketsInternet platformse-commerce