Automatic Metadata Capture and Processing for High-Performance Workflows

📅 2025-06-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In heterogeneous high-performance computing (HPC) environments, workflow metadata collection remains challenging due to fragmentation, poor reusability, and lack of standardization—hindering FAIR (Findable, Accessible, Interoperable, Reusable) compliance and impeding efficient performance analysis. To address this, we propose an automated metadata management framework. It features a lightweight runtime collection mechanism supporting fine-grained metadata capture across heterogeneous workflow systems (e.g., Snakemake, Nextflow); a dual-format unified storage scheme combining JSON Schema (for semantic expressiveness and human readability) and SQLite (for efficient querying); and a performance-aware metadata schema redesign that natively models task dependencies, resource consumption, and temporal behavior. Experimental evaluation demonstrates significant improvements in metadata findability, interoperability, and reusability—enabling reproducible, data-driven workflow performance analysis.

Technology Category

Search and Optimization: Metareasoning and MetaheuristicsData Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, TrustMachine Learning: Hardware-aware ML

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterizationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
Modern workflows run on increasingly heterogeneous computing architectures and with this heterogeneity comes additional complexity. We aim to apply the FAIR principles for research reproducibility by developing software to collect metadata annotations for workflows run on HPC systems. We experiment with two possible formats to uniformly store these metadata, and reorganize the collected metadata to be as easy to use as possible for researchers studying their workflow performance.
Problem

Research questions and friction points this paper is trying to address.

Capture metadata for workflows on heterogeneous architectures
Apply FAIR principles to enhance research reproducibility
Standardize and reorganize metadata for performance analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated metadata capture for HPC workflows
Uniform metadata storage in two formats
Reorganized metadata for performance analysis
🔎 Similar Papers
2024-08-30Scientific DataCitations: 0
💼 Related Jobs
No related jobs found.
P
Polina Shpilker
Tufts University
L
L. Pouchard
Sandia National Laboratories