A Dataset for Named Entity Recognition and Relation Extraction from Art-historical Image Descriptions

📅 2026-02-22
📈 Citations: 0
Influential: 0
📄 PDF

career value

168K/year
🤖 AI Summary
This work addresses the scarcity of fine-grained annotated resources for named entity recognition (NER) and relation extraction (RE) in art history. To bridge this gap, the authors introduce FRAME, a novel dataset comprising descriptions of individual artworks sourced from museum catalogs and auction records. The dataset features a three-tier manual annotation scheme—metadata, content, and coreference layers—covering 37 entity types aligned with Wikidata. Annotations are provided in stand-off format using UIMA XMI CAS, facilitating tasks such as entity linking, knowledge graph construction, and fine-tuning or evaluation of large language models. As the first open-source, fine-grained, multi-layer annotated resource tailored to art historical research, FRAME establishes a benchmark for NER, RE, and few-shot or zero-shot modeling in this domain.

Technology Category

Application Category

📝 Abstract
This paper introduces FRAME (Fine-grained Recognition of Art-historical Metadata and Entities), a manually annotated dataset of art-historical image descriptions for Named Entity Recognition (NER) and Relation Extraction (RE). Descriptions were collected from museum catalogs, auction listings, open-access platforms, and scholarly databases, then filtered to ensure that each text focuses on a single artwork and contains explicit statements about its material, composition, or iconography. FRAME provides stand-off annotations in three layers: a metadata layer for object-level properties, a content layer for depicted subjects and motifs, and a co-reference layer linking repeated mentions. Across layers, entity spans are labeled with 37 types and connected by typed RE links between mentions. Entity types are aligned with Wikidata to support Named Entity Linking (NEL) and downstream knowledge-graph construction. The dataset is released as UIMA XMI Common Analysis Structure (CAS) files with accompanying images and bibliographic metadata, and can be used to benchmark and fine-tune NER and RE systems, including zero- and few-shot setups with Large Language Models (LLMs).
Problem

Research questions and friction points this paper is trying to address.

Named Entity Recognition
Relation Extraction
Art-historical Image Descriptions
Dataset
Entity Annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Named Entity Recognition
Relation Extraction
Art-historical Dataset
Stand-off Annotation
Entity Linking
🔎 Similar Papers
No similar papers found.