Harmonizing Metadata of Language Resources for Enhanced Querying and Accessibility

📅 2025-01-09
📈 Citations: 0
Influential: 0
📄 PDF

career value

152K/year
🤖 AI Summary
Heterogeneous metadata across language resources (LRs) impedes their discoverability and interoperability. Method: This study proposes a unified RDF metadata model integrating the DCAT and META-SHARE ontologies—achieving, for the first time, deep semantic coupling between them. It introduces a novel evaluation paradigm grounded in real user queries (CML) and designs an open-vocabulary-driven, API-based mechanism for machine-readable access. Leveraging the Linked Data technology stack and the Linghub portal, the system supports text search, faceted browsing, and SPARQL querying. Contribution/Results: Empirical evaluation demonstrates significant improvements in LR discoverability, accessibility, and subset extractability; effectively identifies critical metadata heterogeneity issues; and validates that standards-driven ontology integration delivers substantial, measurable gains in LR infrastructure interoperability.

Technology Category

Application Category

📝 Abstract
This paper addresses the harmonization of metadata from diverse repositories of language resources (LRs). Leveraging linked data and RDF techniques, we integrate data from multiple sources into a unified model based on DCAT and META-SHARE OWL ontology. Our methodology supports text-based search, faceted browsing, and advanced SPARQL queries through Linghub, a newly developed portal. Real user queries from the Corpora Mailing List (CML) were evaluated to assess Linghub capability to satisfy actual user needs. Results indicate that while some limitations persist, many user requests can be successfully addressed. The study highlights significant metadata issues and advocates for adherence to open vocabularies and standards to enhance metadata harmonization. This initial research underscores the importance of API-based access to LRs, promoting machine usability and data subset extraction for specific purposes, paving the way for more efficient and standardized LR utilization.
Problem

Research questions and friction points this paper is trying to address.

Language Metadata
Uniformity
Interoperability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linked Data
RDF Technologies
APIs for Metadata Access