Digging Up Citations: FOSSIL, a Dataset and Workflow for Reference Extraction in Law and the Humanities

📅 2026-05-31
📈 Citations: 0
Influential: 0
📄 PDF

career value

165K/year
🤖 AI Summary
Existing citation extraction tools struggle to process footnotes in humanities and legal scholarship due to their embedded placement within main text, inclusion of commentary and cross-references, and highly variable formatting. To address this challenge, this work introduces FOSSIL, the first multilingual open dataset specifically designed for footnote citations, comprising over 7,600 annotated references from 96 scholarly papers. The authors also develop a dedicated PDF-TEI Editor collaborative annotation platform, standardize a seven-annotator labeling protocol, and implement a Grobid-based customized footnote parsing model. This end-to-end pipeline substantially improves performance, increasing the micro F1-score from 0.36 to 0.72 with notable gains in recall, thereby demonstrating the approach’s effectiveness while highlighting remaining challenges in handling cross-references and mixed-content footnotes.
📝 Abstract
Citation extraction tools are designed for the structured end-of-document bibliographies of the natural sciences, but law and humanities scholarship cites references primarily in footnotes, where bibliographic data is interleaved with commentary and cross-references and varies widely across languages and styles. To address the scarcity of suitable gold-standard resources, we present FOSSIL (Footnote-based Open-access SSH Scientific Instance Labels), an openly licensed multilingual dataset of 96 annotated scholarly articles containing over 7,600 footnote-embedded references, together with PDF-TEI Editor (a collaborative web annotation tool), a documented seven-annotator workflow, and a Grobid specialization for footnote-based citations. In end-to-end evaluation, the specialized pipeline nearly doubles extraction quality over default Grobid (micro-F1 from 0.36 to 0.72), driven largely by improved recall, while showing that substantial headroom remains for cross-references and mixed-content footnotes. This extended abstract presents work in progress; annotations of citations segmentation and parsing, and cross-reference resolution are ongoing.
Problem

Research questions and friction points this paper is trying to address.

citation extraction
footnotes
law and humanities
bibliographic data
multilingual
Innovation

Methods, ideas, or system contributions that make the work stand out.

citation extraction
footnote parsing
multilingual dataset
Grobid specialization
humanities scholarship
🔎 Similar Papers
No similar papers found.