How are research data referenced? The use case of the research data repository RADAR

📅 2025-05-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Assessing data reuse and citation practices remains challenging due to fragmented metrics and inconsistent reporting. Method: This study systematically analyzes citation patterns of datasets published in the German RADAR repository, integrating and cross-validating evidence from Google Scholar, DataCite Event Data, and the Data Citation Corpus. It distinguishes “formal citations” (e.g., in reference lists or data availability statements) from “substantive reuse” (i.e., independent scholarly use without author overlap), identified via institutional affiliation matching. Contribution/Results: Among all RADAR datasets, 27.9% received at least one formal citation, of which 21.4% adhered to community citation standards; only 21 datasets (0.5%) evidenced unambiguous external reuse—indicating nascent data reuse maturity. The study establishes a methodological framework for empirically evaluating data impact and advancing robust, interoperable data citation ecosystems.

Technology Category

Data Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB CompletionApplication Domains: Humanities & Computational Social ScienceHumans and AI: Crowd Sourcing and Human Computation

Application Category

Web Mining and Content Analysis: Web data provenance, reliability, and authenticityEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Publishing research data aims to improve the transparency of research results and facilitate the reuse of datasets. In both cases, referencing the datasets that were used is recommended. Research data repositories can support data referencing through various measures and also benefit from it, for example using this information to demonstrate their impact. However, the literature shows that the practice of formally citing research data is not widespread, data metrics are not yet established, and effective incentive structures are lacking. This article examines how often and in what form datasets published via the research data repository RADAR are referenced. For this purpose, the data sources Google Scholar, DataCite Event Data and the Data Citation Corpus were analyzed. The analysis shows that 27.9 % of the datasets in the repository were referenced at least once. 21.4 % of these references were (also) present in the reference lists and are therefore considered data citations. Datasets were referenced often in data availability statements. A comparison of the three data sources showed that there was little overlap in the coverage of references. In most cases (75.8 %), data and referencing objects were published in the same year. Two definition approaches were considered to investigate data reuse. 118 RADAR datasets were referenced more than once. Only 21 references had no overlaps in the authorship information -- these datasets were referenced by researchers that were not involved in data collection.
Problem

Research questions and friction points this paper is trying to address.

Examining dataset referencing frequency and forms in RADAR
Assessing data citation practices and their impact metrics
Investigating data reuse through authorship and reference patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzed Google Scholar, DataCite, Citation Corpus
Examined RADAR dataset referencing frequency
Compared authorship for data reuse cases
🔎 Similar Papers
No similar papers found.
D
Dorothea Strecker
Berlin School of Library and Information Science, Humboldt-Universität zu Berlin
K
Kerstin Soltau
FIZ Karlsruhe - Leibniz Institute for Information Infrastructure
F
Felix Bach
FIZ Karlsruhe - Leibniz Institute for Information Infrastructure