reference verification

Designs, builds, or evaluates methods and tools that verify bibliographic citations by resolving entries to authoritative records, checking that referenced works exist, and validating citation metadata (authors, titles, publication details). Work includes detecting non‑existent or duplicate reference records, identifying substantial author‑list or other metadata mismatches, and cross‑checking entries across multiple sources to flag inconsistencies.

referenceverification

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses persistent issues in academic manuscripts—such as erroneous citation identifiers, missing metadata, misattributed authorship, and confusion between preprints and published versions—exacerbated by the propensity of large language models to generate citation hallucinations. To mitigate these challenges, we propose a TypeScript-based Model Context Protocol (MCP) server that integrates automated citation validation into intelligent scholarly editing workflows for the first time. Our system employs a manifestation-aware matching mechanism and policy-gated rewriting strategies, harmonizing data from multiple sources including PubMed, Crossref, arXiv, and Semantic Scholar. It supports structured parsing of diverse file formats and multi-round retrieval to generate precise correction suggestions. The prototype has been rigorously evaluated across 47 test cases covering repair actions, exception handling, and protocol compliance, demonstrating robust defense against both conventional citation errors and LLM-induced hallucinations.

bibliographic errorscitation hallucinationLLM-induced errors

To address the lack of systematic quality assurance for bibliographic and citation data in the OpenCitations infrastructure, this paper designs and implements an interpretable validation and dynamic quality monitoring framework tailored to the OpenCitations Data Model (OCDM). Methodologically, it integrates a customizable rule engine, SPARQL-based consistency checking, semantic constraint validation, and an incremental quality dashboard, enabling error attribution analysis and quantitative assessment. Key contributions include: (1) the first interpretable validation tool specifically designed for OCDM; and (2) a novel dynamic, sustainable quality tracking mechanism. Experimental evaluation demonstrates that the framework accurately identifies structural and semantic defects in the Matilda dataset and detects, localizes, and quantifies persistent issues—including duplication, incompleteness, and inconsistency—in OpenCitations Meta. The approach significantly enhances data reliability and fills a critical gap in systematic quality assurance for open citation data.

Ensuring quality in bibliographic and citation datasetsMonitoring data quality in open research infrastructuresValidating metadata and citations from diverse sources

Over 698 million references in Crossref lack DOIs, severely impeding the structural construction of citation networks. This paper proposes a systematic approach integrating heuristic rules, metadata matching, and fuzzy text parsing to accurately align unstructured references to target文献 entities in OpenCitations Meta. Methodologically, it combines an interpretable rule engine with semantic matching strategies—specifically designed to enhance linkage accuracy under sparse or inconsistent metadata conditions. Contributions include: (1) a novel hybrid alignment framework balancing precision and robustness; (2) a manually curated gold standard and a validated Crossref subset for rigorous evaluation. Experimental results demonstrate high precision, substantially expanding both the coverage breadth and linkage quality of the open citation network. The method provides a reproducible, scalable technical pathway for large-scale reference DOI enrichment.

Addressing metadata inconsistencies in Crossref references using heuristic toolsCreating formal citation links when references lack specified DOIsMatching bibliographic references with incomplete metadata to existing entities

Latest Papers

What's happening recently
View more

Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.

academic writinghallucinated citationsreference reliability

This work addresses the growing threat of hallucinated citations—fabricated references generated by large language models—that have infiltrated top-tier academic venues such as ICLR, ICML, NeurIPS, and USENIX Security, undermining scholarly credibility. To combat this, the authors introduce RefChecker, the first scalable citation verification pipeline tailored for large-scale analysis of conference papers. RefChecker integrates multi-source academic database matching with web-based re-verification to efficiently assess citation authenticity under conservative criteria. The study presents the first systematic quantification of citation hallucinations under a rigorous definition, revealing that approximately 5% of NeurIPS and USENIX Security papers—including some award-winning works—contain at least two hallucinated references. Notably, the approach achieves audit costs as low as $0.04 per paper, demonstrating that automated, low-cost, and reproducible large-scale citation auditing is both feasible and practical.

academic integritybibliographic verificationcitation hallucination

Hot Scholars

MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
AD

Amit Dhurandhar

Principal Research Scientist, IBM
artificial intelligencemachine learningdata mining
JJ

Junfeng Jiao

Associate Professor, Urban Information Lab, Texas Smart City, NSF NRT AI, UT Austin
AISmart CityUrban Informatics
SC

Stephen Casper

PhD student, MIT
AI safetyAI responsibilityred-teamingrobustness
TP

Tao Peng

吉林大学
natural language processingknowledge graph