annotate data semantically

Design, apply, and manage machine-readable semantic annotations and ontology-linked metadata for datasets, mapping fields and values to controlled concepts, relationships, and provenance records. Build annotation schemas and tooling to capture contextual information, enable dataset interoperability, reuse, and automated semantic processing.

annotatedatasemantically

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A metadata model for profiling multidimensional sources in data ecosystems

Mar 20, 2025
CD
C. Diamantini
🏛️ Università Politecnica delle Marche

To address insufficient semantic description of multidimensional aggregate/summary data, poor adaptability of metadata standards, and cross-source interoperability challenges in big data environments, this paper proposes a multidimensional data source profiling metadata model tailored for data ecosystems. Built upon RDF, the model is the first to support extensible semantic modeling of both aggregate and summary multidimensional data, enabling semantic alignment of dimensions and measures with reference knowledge graphs. It integrates multi-granularity metadata profiles—spanning source-level, attribute-level, and value-distribution characteristics. The model ensures flexible extensibility and cross-source interoperability. Experimental results demonstrate that profile generation time scales linearly with data cardinality, confirming its engineering practicality and predictable performance.

Challenges in managing diverse data formats in Big Data ecosystems.Inadequate metadata vocabularies for aggregated or summary data.Need for efficient metadata model for multidimensional data profiling.

Knowledge engineering for open science: Building and deploying knowledge bases for metadata standards

Jul 30, 2025
MA
Mark A. Musen
🏛️ Stanford Center for Biomedical Informatics Research | Stanford University School of Medicine

To address insufficient standardization, poor domain adaptability, and practical challenges in implementing FAIR principles for research data metadata in open science, this paper proposes a knowledge engineering–based metadata templating approach. It formalizes domain-specific metadata standards as reusable, logically inferable knowledge templates; designs lightweight, dual-mode (Web form and spreadsheet) acquisition interfaces with real-time validation; and enables cross-platform intelligent integration via declarative knowledge representation. The method has been adopted by multiple international scientific consortia to establish standardized metadata frameworks and has driven the development of over ten data annotation systems. Empirical outcomes demonstrate significant improvements in metadata syntactic and semantic consistency, domain alignment, and cross-platform interoperability. By embedding domain knowledge into machine-processable templates, the approach delivers a scalable, maintainable, knowledge-driven paradigm for FAIR data infrastructure.

Building knowledge bases for FAIR metadata standardsCreating templates to standardize scientific metadataDeploying metadata standards in intelligent systems

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

Apr 08, 2024
SS
Sowmya S. Sundaram
🏛️ Stanford University

This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.

Enhance metadata standards adherenceImprove metadata curation automationIntegrate structured knowledge with LLMs

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing ontology documentation tools in supporting modular modeling and human readability, particularly in handling cross-module entities and annotations. To overcome these challenges, the authors refactor and extend the LODE framework by introducing a modular Reader-Model-Viewer architecture that decouples parsing, modeling, and rendering components. Implemented as a web service, the new framework provides enhanced capabilities for generating OWL ontology documentation, featuring dedicated entity pages, RDF provenance tracking, and Markdown-based rendering. These improvements significantly increase the intelligibility and reusability of modular scientific knowledge graph ontologies. The framework has been successfully applied to the documentation of the SKG-O ontology, demonstrating its practical utility and effectiveness.

human-readable documentationmodular ontologiesontology documentation

This study investigates whether structured semantic metadata—such as schema.org—remains essential for intelligent agents to achieve reliable and executable data retrieval in the era of large language models (LLMs). By constructing an LLM-as-a-judge evaluation framework, the authors systematically compare semantic-aware agents leveraging such metadata against baseline agents relying solely on open web content, assessing their performance under the FAIR (Findable, Accessible, Interoperable, Reusable) principles. Experimental results demonstrate that semantic agents achieve a 65.7% improvement in overall precision for retrieving FAIR-compliant datasets and a 46.6% gain in identifying results with machine-readable download links. This work provides the first quantitative evidence of the “last-mile utility” of the semantic web ecosystem for executable tasks, underscoring the enduring value of structured metadata even in the age of LLMs.

agentic data retrievalFAIR principlesLarge Language Models

This study addresses the challenge of low-quality metadata that hinders dataset discoverability and reuse, particularly in the context of large language model (LLM)-generated descriptions lacking empirical guidance on context selection and its impact on quality. Building a literature-based framework for description quality assessment, the authors conduct systematic ablation experiments across 252 real-world CSV datasets. They uncover a previously unreported “table-structure penalty” phenomenon: relying solely on table structure significantly degrades narrative quality. While representative data samples aid semantic grounding, they do not improve overall human-rated quality. The work further reveals that different LLMs exhibit consistent descriptive styles. Through LLM-as-a-judge evaluations, semantic attribute analysis, and large-scale experimentation, the study offers key recommendations for LLM-assisted data publishing: concise, relevant context yields better results than redundant input, and table structure should be used cautiously as a basis for generation.

context ablationdata reusedataset description

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

Hot Scholars

YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
YC

Yubo Chen

Institute of Automation, Chinese Academy of Sciences
Natural Language ProcessingInformation ExtractionEvent ExtractionLarge Language Model
EB

Eva Blomqvist

Professor in Computer Science, Linköping University
Semantic WebArtificial IntelligenceKnowledge GraphsOntology Engineering
SA

Sören Auer

Leibniz University of Hannover, Leibniz TIB, L3S Research Center
Neurosymbolic AIKnowledge GraphsWeb ScienceDigital Libraries
JK

Jacques Klein

University of Luxembourg / SnT
Computer ScienceSoftware EngineeringAndroid SecuritySoftware Security