🤖 AI Summary
This study addresses persistent challenges in open data publishing within industry–academia–government collaboration, including inefficient data management, barriers to data reuse, weak licensing awareness, and insufficient integration of real and synthetic data. Drawing on in-depth analysis of 13 European collaborative project datasets, statistical examination of metadata from 281,000 datasets on Zenodo, and complementary surveys and inductive reasoning, the study reveals three key empirical findings: (1) data collection planning plays a critical, previously underrecognized role; (2) script documentation is extremely rare (only 2.4% of datasets); and (3) licensing practices are widespread but largely noncompliant. It further provides robust evidence that hybrid real-synthetic or simulation-based datasets hold substantial scientific value. Based on these insights, the study proposes an actionable data management framework and concrete standardization recommendations—aimed at enhancing cross-sectoral data reusability, regulatory compliance, and the maturity of open science practices.
📝 Abstract
Effective data management and sharing are critical success factors in industry-academia collaboration. This paper explores the motivations and lessons learned from publishing open data sets in such collaborations. Through a survey of participants in a European research project that published 13 data sets, and an analysis of metadata from almost 281 thousand datasets in Zenodo, we collected qualitative and quantitative results on motivations, achievements, research questions, licences and file types. Through inductive reasoning and statistical analysis we found that planning the data collection is essential, and that only few datasets (2.4%) had accompanying scripts for improved reuse. We also found that authors are not well aware of the importance of licences or which licence to choose. Finally, we found that data with a synthetic origin, collected with simulations and potentially mixed with real measurements, can be very meaningful, as predicted by Gartner and illustrated by many datasets collected in our research project.