๐ค AI Summary
This study addresses the longstanding challenge of effectively measuring the real-world impact of research data reuse in open science. For the first time, it leverages large language models (LLMs) and generative artificial intelligence to conduct large-scale analysis of scholarly texts, automatically identifying instances where published works reuse existing datasets. The proposed approach achieves substantially higher detection accuracy compared to conventional methods, revealing a data reuse rate of 43%โfar exceeding estimates derived from traditional bibliometric indicators. These findings demonstrate that current metrics systematically underestimate the benefits of data sharing and establish a novel paradigm for evaluating the impact of open science initiatives.
๐ Abstract
Numerous metascience studies and other initiatives have begun to monitor the prevalence of open science practices when it is more important to understand the 'downstream' effects or impacts of open science. PLOS and DataSeer have developed a new LLM-based indicator to measure an important effect of open science: the reuse of research data. Our results show a data reuse rate of 43%, which is higher than established bibliometric techniques. We show that data reuse can be measured at scale using LLMs and generative artificial intelligence. The positive effects of research data sharing and reuse may currently be underestimated.