manage experimental data

Designs and implements systems, processes, and artifacts to collect, store, organize, and provision experimental data across multiple sources, including instrumenting collection pipelines and assembling multi-source data cohorts. Builds and enforces metadata and format standards, provenance tracking, and data-quality monitoring so datasets are reproducible, discoverable, and reliable for downstream analysis.

manageexperimentaldata

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of reproducibility in actively developed experimental projects, which often suffer from unstructured data management and are overlooked by conventional data management plans. We propose a lightweight, domain-agnostic framework built upon the Sacred experiment tracking model that, from the project’s inception, systematically organizes parameters, metadata, metric trajectories, and associated files. Small-scale data are stored in a NoSQL database, while large files are linked via unique identifiers to dedicated storage systems. The framework seamlessly integrates into existing research workflows, supports both local deployment and public release, and uniquely targets the dynamic exploration phase of research. By doing so, it establishes a practical bridge from early-stage experimentation to FAIR-compliant data sharing, significantly enhancing collaborative efficiency and scientific reproducibility without compromising flexibility or scalability.

collaborative researchexperimental data managementFAIR data

The hunt for research data: Development of an open-source workflow for tracking institutionally-affiliated research data publications

Jul 01, 2025
BM
Bryan M. Gee
🏛️ University of Texas Libraries | The University of Texas at Austin

Institutions face significant challenges in systematically tracking their affiliated research data publications, primarily due to the widespread absence, inconsistency, or non-standardization of institutional attribution metadata (e.g., missing or ambiguous institutional names, lack of persistent identifiers such as DOIs) in existing data repositories. Method: We propose the first open-source, institution-centric workflow for tracking data publications, integrating over 70 open APIs and employing multi-source metadata harvesting, normalization, and automated aggregation—thereby reducing reliance on explicit attribution signals like DOIs or manually curated affiliations. Contribution/Results: The workflow enables efficient discovery and consolidation of over 4,000 cross-platform datasets. Evaluation demonstrates substantial improvements in institutional data discoverability, coverage breadth, and retrieval efficiency. This work delivers a reusable technical infrastructure to support research administration and data governance at institutional and systemic levels.

Address challenges in discovering institution-affiliated datasetsDevelop open-source workflow for tracking institutional research dataImprove metadata standardization for comprehensive data retrieval

Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.

Addressing spreadsheet limitations for consistent experiment-related metadata annotationEnsuring metadata standards compliance in spreadsheet-based scientific data entryProviding quality control for biomedical metadata collection using spreadsheets

Automatic Metadata Capture and Processing for High-Performance Workflows

Jun 18, 2025
PS
Polina Shpilker
🏛️ Tufts University | Sandia National Laboratories

In heterogeneous high-performance computing (HPC) environments, workflow metadata collection remains challenging due to fragmentation, poor reusability, and lack of standardization—hindering FAIR (Findable, Accessible, Interoperable, Reusable) compliance and impeding efficient performance analysis. To address this, we propose an automated metadata management framework. It features a lightweight runtime collection mechanism supporting fine-grained metadata capture across heterogeneous workflow systems (e.g., Snakemake, Nextflow); a dual-format unified storage scheme combining JSON Schema (for semantic expressiveness and human readability) and SQLite (for efficient querying); and a performance-aware metadata schema redesign that natively models task dependencies, resource consumption, and temporal behavior. Experimental evaluation demonstrates significant improvements in metadata findability, interoperability, and reusability—enabling reproducible, data-driven workflow performance analysis.

Apply FAIR principles to enhance research reproducibilityCapture metadata for workflows on heterogeneous architecturesStandardize and reorganize metadata for performance analysis

Latest Papers

What's happening recently
View more

This study investigates how domain-specific metadata schemas can be effectively integrated with the generic DataCite schema to enhance metadata quality and interoperability in research data repositories. Through structural comparisons, cross-schema mapping analyses, and workflow evaluations of metadata records from eight repositories in the earth and social sciences, the research reveals how disciplinary characteristics influence the completeness of DataCite records. Findings indicate that discrepancies between schemas stem primarily from differing modeling philosophies rather than expressive capacity. While optimized cross-schema mappings significantly improve metadata quality, the diversity of repository workflows also critically affects record completeness. Building on these insights, the study proposes a strategy that leverages the complementary strengths of domain-specific and generic schemas, offering practical guidance for fostering interdisciplinary data sharing.

DataCitedisciplinary metadatametadata interoperability

Addressing challenges in FAIR principle implementation—including fragmented data and code lifecycles, lack of executable environments, and high technical barriers—this study proposes a unified open-science platform. The platform uniquely integrates version control, containerized computational environments, and modular project scaffolding to support end-to-end reproducible research, from grant proposal to publication. It interoperates with mainstream scientific toolchains, supports deployment on both local workstations and institutional servers, and provides a lightweight graphical user interface. Empirical validation demonstrates successful re-execution of over a dozen interdisciplinary studies published more than ten years ago, confirming the platform’s robust long-term reproducibility, cross-platform compatibility, and seamless execution across diverse domains. By significantly lowering technical adoption barriers for researchers, the platform enables practical integration of FAIR principles and reproducibility practices into routine scientific workflows.

Bridges disconnected data and code life cyclesEnables FAIR workflows without manual setupUnifies data and software lifecycles for reproducibility

Experiversum: an Ecosystem for Curating and Enhancing Data-Driven Experimental Science

Sep 30, 2025
GV
Genoveva Vargas-Solar
🏛️ CNRS | Univ Lyon | INSA Lyon | UCBL | LIRIS | Federal University Rio Grande do Norte | Université Lumière Lyon 2 | Université Claude Bernard Lyon 1 | Federal University of Parana | Universidad de las República | Fundación Universidad de las Américas Puebla

In data-driven experimental science, challenges persist regarding disorganized data curation, insufficient documentation, and poor reproducibility. To address these, this paper proposes a lakehouse-based scientific data governance ecosystem. The system integrates automated metadata capture, joint data-and-code versioning, and collaborative, auditable decision logging, establishing an iterative closed loop spanning data acquisition, processing, analysis, and interpretation—thereby harmonizing exploratory research with reproducibility requirements. Crucially, it embeds dynamic, context-aware governance directly into the research workflow. Its cross-disciplinary applicability is empirically validated across heterogeneous domains: Earth science, life science, and political science. Results demonstrate significant improvements in process transparency, auditability, and result interpretability. The ecosystem provides a scalable, verifiable infrastructure for multi-disciplinary collaborative research.

Bridges exploratory and reproducible research across disciplinesEnables structured research through iterative data cyclesFacilitates curation and reproducibility of exploratory experiments