Score
Designs and builds pipelines and parsers that extract citation metadata, identifiers, and reference structures from documents and match those citations to primary records or bibliographic entries. Designs and analyzes verification systems that retrieve and inspect cited sources to confirm the cited content supports specific claims, establish claim–evidence links, and flag unverifiable, missing, or mismatched references.
The evidence-based text generation (EBTG) field suffers from terminological ambiguity, fragmented evaluation practices, and the absence of standardized benchmarks. Method: We conduct a systematic review of 134 papers and propose the first unified taxonomy covering three core operations—citation, attribution, and quotation—structured along seven dimensions (e.g., evidence source, granularity, explicitness). Concurrently, we design a multidimensional evaluation framework that integrates and standardizes 300 metrics, and introduce the first reproducible, comparable EBTG benchmark. Contribution/Results: Our work resolves long-standing fragmentation in EBTG research, establishes a rigorous foundation for assessing output traceability, verifiability, and trustworthiness, and provides both theoretical guidance and practical tools to advance reliable large language model development.
This work addresses the critical challenge of hallucinated citations in scientific texts generated by large language models, which pose a serious threat to academic integrity and are difficult to detect manually. To this end, the authors introduce the first benchmark and verification framework specifically designed for detecting fabricated references in scientific writing. The framework features a novel unified metric for evaluating citation faithfulness and evidence alignment, supported by a large-scale, cross-domain dataset rigorously validated by human annotators. Methodologically, it proposes an interpretable and scalable multi-agent verification pipeline that integrates claim extraction, evidence retrieval, passage matching, and reasoning calibration to enable end-to-end auditing of citation authenticity. Experimental results demonstrate that the proposed approach significantly outperforms existing methods in both accuracy and interpretability, effectively identifying citation errors produced by state-of-the-art large language models and offering a reliable tool for scientific publishing integrity.
Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.
Over 698 million references in Crossref lack DOIs, severely impeding the structural construction of citation networks. This paper proposes a systematic approach integrating heuristic rules, metadata matching, and fuzzy text parsing to accurately align unstructured references to target文献 entities in OpenCitations Meta. Methodologically, it combines an interpretable rule engine with semantic matching strategies—specifically designed to enhance linkage accuracy under sparse or inconsistent metadata conditions. Contributions include: (1) a novel hybrid alignment framework balancing precision and robustness; (2) a manually curated gold standard and a validated Crossref subset for rigorous evaluation. Experimental results demonstrate high precision, substantially expanding both the coverage breadth and linkage quality of the open citation network. The method provides a reproducible, scalable technical pathway for large-scale reference DOI enrichment.
This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.
Existing citation extraction tools struggle to process footnotes in humanities and legal scholarship due to their embedded placement within main text, inclusion of commentary and cross-references, and highly variable formatting. To address this challenge, this work introduces FOSSIL, the first multilingual open dataset specifically designed for footnote citations, comprising over 7,600 annotated references from 96 scholarly papers. The authors also develop a dedicated PDF-TEI Editor collaborative annotation platform, standardize a seven-annotator labeling protocol, and implement a Grobid-based customized footnote parsing model. This end-to-end pipeline substantially improves performance, increasing the micro F1-score from 0.36 to 0.72 with notable gains in recall, thereby demonstrating the approach’s effectiveness while highlighting remaining challenges in handling cross-references and mixed-content footnotes.
This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.