🤖 AI Summary
In heterogeneous high-performance computing (HPC) environments, workflow metadata collection remains challenging due to fragmentation, poor reusability, and lack of standardization—hindering FAIR (Findable, Accessible, Interoperable, Reusable) compliance and impeding efficient performance analysis. To address this, we propose an automated metadata management framework. It features a lightweight runtime collection mechanism supporting fine-grained metadata capture across heterogeneous workflow systems (e.g., Snakemake, Nextflow); a dual-format unified storage scheme combining JSON Schema (for semantic expressiveness and human readability) and SQLite (for efficient querying); and a performance-aware metadata schema redesign that natively models task dependencies, resource consumption, and temporal behavior. Experimental evaluation demonstrates significant improvements in metadata findability, interoperability, and reusability—enabling reproducible, data-driven workflow performance analysis.
📝 Abstract
Modern workflows run on increasingly heterogeneous computing architectures and with this heterogeneity comes additional complexity. We aim to apply the FAIR principles for research reproducibility by developing software to collect metadata annotations for workflows run on HPC systems. We experiment with two possible formats to uniformly store these metadata, and reorganize the collected metadata to be as easy to use as possible for researchers studying their workflow performance.