dataset reporting standards

Design, build, and evaluate dataset reporting standards and metadata systems—including schemas, documentation templates, provenance and uncertainty fields, and validation rules—that enable consistent labeling, interoperability, and cross-study comparability. Implement tools and processes for metadata extraction, parsing, curation, linking, integration, management, quality auditing, filtering, and feature engineering to support dataset discovery, reuse, and downstream analysis.

datasetreportingstandards

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$185K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Knowledge engineering for open science: Building and deploying knowledge bases for metadata standards

Jul 30, 2025
MA
Mark A. Musen
🏛️ Stanford Center for Biomedical Informatics Research | Stanford University School of Medicine

To address insufficient standardization, poor domain adaptability, and practical challenges in implementing FAIR principles for research data metadata in open science, this paper proposes a knowledge engineering–based metadata templating approach. It formalizes domain-specific metadata standards as reusable, logically inferable knowledge templates; designs lightweight, dual-mode (Web form and spreadsheet) acquisition interfaces with real-time validation; and enables cross-platform intelligent integration via declarative knowledge representation. The method has been adopted by multiple international scientific consortia to establish standardized metadata frameworks and has driven the development of over ten data annotation systems. Empirical outcomes demonstrate significant improvements in metadata syntactic and semantic consistency, domain alignment, and cross-platform interoperability. By embedding domain knowledge into machine-processable templates, the approach delivers a scalable, maintainable, knowledge-driven paradigm for FAIR data infrastructure.

Building knowledge bases for FAIR metadata standardsCreating templates to standardize scientific metadataDeploying metadata standards in intelligent systems

10 Simple Rules for Improving Your Standardized Fields and Terms

Oct 21, 2025
RC
Rhiannon Cameron
🏛️ Simon Fraser University

Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.

Addressing challenges in standardizing research metadata vocabulariesOffering practical rules for FAIR-compliant metadata designProviding strategies to improve data findability and reusability

Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.

Addressing spreadsheet limitations for consistent experiment-related metadata annotationEnsuring metadata standards compliance in spreadsheet-based scientific data entryProviding quality control for biomedical metadata collection using spreadsheets

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

Apr 08, 2024
SS
Sowmya S. Sundaram
🏛️ Stanford University

This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.

Enhance metadata standards adherenceImprove metadata curation automationIntegrate structured knowledge with LLMs

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Latest Papers

What's happening recently
View more

This study investigates how domain-specific metadata schemas can be effectively integrated with the generic DataCite schema to enhance metadata quality and interoperability in research data repositories. Through structural comparisons, cross-schema mapping analyses, and workflow evaluations of metadata records from eight repositories in the earth and social sciences, the research reveals how disciplinary characteristics influence the completeness of DataCite records. Findings indicate that discrepancies between schemas stem primarily from differing modeling philosophies rather than expressive capacity. While optimized cross-schema mappings significantly improve metadata quality, the diversity of repository workflows also critically affects record completeness. Building on these insights, the study proposes a strategy that leverages the complementary strengths of domain-specific and generic schemas, offering practical guidance for fostering interdisciplinary data sharing.

DataCitedisciplinary metadatametadata interoperability

This study addresses the challenge of low-quality metadata that hinders dataset discoverability and reuse, particularly in the context of large language model (LLM)-generated descriptions lacking empirical guidance on context selection and its impact on quality. Building a literature-based framework for description quality assessment, the authors conduct systematic ablation experiments across 252 real-world CSV datasets. They uncover a previously unreported “table-structure penalty” phenomenon: relying solely on table structure significantly degrades narrative quality. While representative data samples aid semantic grounding, they do not improve overall human-rated quality. The work further reveals that different LLMs exhibit consistent descriptive styles. Through LLM-as-a-judge evaluations, semantic attribute analysis, and large-scale experimentation, the study offers key recommendations for LLM-assisted data publishing: concise, relevant context yields better results than redundant input, and table structure should be used cautiously as a basis for generation.

context ablationdata reusedataset description

This study addresses the lack of empirical evaluation regarding whether existing dataset documentation frameworks effectively foster developer reflectivity. Combining mixed-methods thematic analysis with corpus-assisted discourse analysis, the research systematically examines how prevailing documentation frameworks—and their real-world instantiations—cover core dimensions of reflectivity. The findings reveal, for the first time, that current frameworks consistently overlook critical reflective themes. Building on this insight, the authors develop a reflectivity-oriented coding manual and propose an enhanced datasheet template incorporating targeted prompts to elicit deeper reflection. This work offers actionable strategies and practical tools to strengthen the reflective capacity of dataset documentation practices.

dataset developmentdatasheetsFAcCT

Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.

heterogeneous datareproducibilityschema

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
SP

Silvio Peroni

University of Bologna
Semantic PublishingSemantic WebOpen ScienceScience of Science
PM

Philipp Mayr

GESIS - Leibniz Institute for the Social Sciences
Interactive Information RetrievalInformetricsDigital librariesInformation Seeking
IH

Ivan Heibi

University of Bologna
Semantic PublishingSemantic WebData VisualisationWeb technologies
AJ

Alexis Joly

Research Director, Inria, Montpellier University, LIRMM
machine learningbiodiversityinformation retrievalplant identification