๐ค AI Summary
This study addresses the persistent challenge of code reusability in computational linguistics (*CL) research by presenting the first quantitative assessment of the long-term availability of GitHub repositories. Employing bibliometric analysis, web crawling, and statistical methods, it systematically examines the survival trends of repositories accompanying papers from top-tier venues such as ACL over the past decade. The findings reveal that repository failure rates have not declined but instead increased in recent years, with a proliferation of empty placeholder repositories emerging as a novel risk. In contrast, non-GitHub platforms demonstrate comparatively greater stability. By exposing the severe attrition of open-source research artifacts, this work provides critical empirical evidence to inform strategies for enhancing academic reproducibility within the *CL community.
๐ Abstract
Source code and data published at computational linguistics (*CL) venues are increasingly being shared via GitHub. While this generally is a favourable development for the accessibility and potential reusability of research artifacts in natural language processing (NLP), the long-term availability of such repositories has not been evaluated. In this squib, we discuss the availability of repositories linked in papers published in the Computational Linguistics (CL) journal as well as at ACL and its co-located events over the past ten years. Contrary to our expectations, we find that GitHub repositories linked in more recent ACL publications are unavailable at similar rates as in older publications, in parts due to an increase in empty and placeholder repositories. Similar trends hold for other *CL venues, but not for platforms other than GitHub.