Benchmarking Automated Knowledge Graph Construction from Semi-Structured Data

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对半结构化数据构建知识图谱的问题,提出了一种结合语法有效性、语义准确性等六个质量维度的评估基准和流程。
📝 Abstract
Knowledge Graphs (KGs) play an increasingly important role in numerous applications ranging from traditional knowledge representation to serving as memory for LLMs to support downstream tasks. However, their construction is labor-intensive; thus, in recent years, numerous approaches for automizing this process have been proposed. Compared to the popularity of construction approaches that focus on textual input data, methods for semi-structured inputs remain underrepresented and as a result, no comprehensive benchmark and evaluation suite exists to judge the quality of mapping predictions and generated KGs. This is problematic, as a KG's quality has direct influence on the downstream applications it supports and thus, strong evaluation mechanisms for their construction are urgently needed. In this work, we thus focus on the evaluation of KG construction from semi-structured data and present a benchmark and evaluation pipeline for KG construction that combines the quality dimensions (1) syntactic validity, (2) semantic accuracy, (3) consistency, (4) conciseness, (5) completeness, and (6) pragmatic quality measured on a KG's ability to provide answers to competency questions. This work contributes a realistic task definition, extends current state of the art evaluation frameworks, allows evaluation of systems that predict mappings and RDF data alike, and combines evaluation of both KG construction and downstream usage. We provide a comprehensive metrics suite, provide ten expert-curated datasets from seven domains, and showcase evaluation using two reference systems.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Graphs
semi-structured data
benchmarking
evaluation
quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

benchmark
semi-structured data
knowledge graph construction
evaluation pipeline
quality dimensions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tarek Al Mustafa
German Centre for Integrative Biodiversity Research Halle-Jena-Leipzig – iDiv, Puschstr. 4, 04103 Leipzig, Germany
Birgitta König-Ries
Birgitta König-Ries
Heinz-Nixdorf Endowed Chair for Distributed Information Systems, University of Jena
Distributed Information SystemsBiodiversity Informatics(Semantic) WebPortal Technology