design structured metadata

Designs and documents structured metadata schemas and information models that specify named fields, data types, allowed values/value sets, validation rules, required vs optional status, and logical grouping of fields into blocks. Builds mappings and notation to ensure interoperability with external standards (e.g., ISIC‑compatible models) and to capture contextual attributes and provenance needed to interpret records.

designstructuredmetadata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$198K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge in attributed graph schema design of whether repeatedly occurring descriptive attributes should be embedded within nodes or externalized as reusable metadata. Building upon Fifth Normal Form (5NF), the authors propose a principled decision framework that systematically identifies metadata candidates based on semantic criteria rather than mere repetition frequency. The approach classifies attributes into characteristic nodes, embedded properties, or borderline cases using five key principles: cross-element occurrence frequency, conceptual independence, lossless externalizability, reuse potential, and governance relevance. Empirical validation through a library domain case study and an entity classification task demonstrates that repetition alone is insufficient for externalization decisions—semantic judgment is essential. The proposed method significantly enhances the accuracy, consistency, and reusability of metadata modeling in graph-based systems.

embedded propertiesmetadataproperty graph schemas

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Although scientific data increasingly adhere to the FAIR principles and employ standardized identifiers, practical interoperability remains hindered by heterogeneity in identifier systems and data models. This work proposes and implements two synergistic tools—Babel and ORION—to bridge this gap. Babel constructs clusters of equivalent identifiers through mapping-based clustering and exposes them via a high-performance quantitative API, while ORION standardizes heterogeneous knowledge bases by aligning them to a community-governed common data model. Together, they systematically address the longstanding disconnect between the FAIR “Interoperable” principle and its real-world implementation. The integration of these tools has enabled the construction of a fully interoperable knowledge base, substantially enhancing cross-resource data integration and query capabilities. The resulting framework is publicly available.

Data ModelsFAIRIdentifier Schemas

This study addresses the lack of systematic guidance on contextualization strategies for large language model (LLM) agents operating in structured data environments, particularly concerning effectiveness and efficiency across multi-file, large-scale schemas. Using SQL generation as a proxy task, the work presents the first systematic evaluation of eleven models across four context formats—YAML, Markdown, JSON, and TOON—at schema scales ranging from 10 to 10,000 tables. The findings reveal that model capability tiers critically determine optimal context architecture: tailored strategies significantly improve performance, with state-of-the-art models gaining 2.7% accuracy under native file-based contexts, while open-source models average a 7.7% decline. Moreover, native file-based agents scale efficiently to ten-thousand-table schemas while maintaining high navigation accuracy.

context engineeringfile-native systemsLLM agents

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

Apr 08, 2024
SS
Sowmya S. Sundaram
🏛️ Stanford University

This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.

Enhance metadata standards adherenceImprove metadata curation automationIntegrate structured knowledge with LLMs

Latest Papers

What's happening recently
View more

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.

graph queriesoutput schemaproperty graphs

This study addresses the challenges posed by the high heterogeneity of healthcare data and the lack of effective metadata management, which often degrade conventional data lakes into “data swamps,” impeding data interoperability and machine learning (ML) readiness. To overcome these limitations, the authors propose a dual-hybrid semantic data lake architecture that synergistically integrates the dynamic modeling capabilities of knowledge graphs with the metadata generation power of large language models (LLMs). A human-in-the-loop validation mechanism is incorporated to enable automated metadata annotation and high-level semantic alignment. This approach establishes, for the first time, semantic linkages within a data lake explicitly oriented toward ML operability, substantially enhancing the discoverability and computability of heterogeneous medical data while supporting intelligent recommendation of suitable ML methods.

data lakeheterogeneous datainteroperability

This study addresses the lack of effective validation mechanisms for cross-ontology semantic mappings by proposing a verification framework grounded in ontological metaphysical commitments. The approach uniquely incorporates metaphysical choices as constraints, employing cardinality restrictions to formally specify mappings between distinct ontologies—such as IES and BFO—and leverages SPARQL queries for automated validation. By integrating ontological analysis, modeling of metaphysical commitments, and precise definition of cardinality constraints, the work establishes an operational validation pipeline demonstrated in an IES–BFO mapping case study. This framework significantly enhances the logical consistency and semantic quality of ontology alignments, thereby offering both theoretical grounding and a practical pathway for achieving semantic interoperability across heterogeneous ontologies.

cardinality constraintsfoundation ontologiesmetaphysical commitments

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

Hot Scholars

RM

Rashid Mushkani

University of Montreal I Mila
Public (Space & Life)Sociotechnical AIUrban AnalyticsCommunity-Centered AI
TK

Tri Kurniawan Wijaya

Huawei Research
Recommender SystemsDeep LearningDemand ResponseSmart Grid
XS

Xinyang Shao

Machine Learning Engineer in Huawei Ireland Research Centre
Recommender System
ZZ

Zhiming Zhao

Associate professor - University of Amsterdam
Cloud computingBig data managementSoftware defined networkingMulti Agent systems
NS

Nafiseh Soveizi

Postdoctoral Researcher at University of Amsterdam
Anomaly detectionPredictive MonitoringWorkflow AdaptationSecurity