large-scale citation analysis

Design and implement systems and pipelines that parse, normalize, and link large collections of citation and bibliographic records into structured metadata (including name disambiguation and citation parsing). Analyze and annotate citation errors and other metadata defects at scale to quantify error rates and completeness across venues and to track temporal trends such as name-change or update coverage.

large-scalecitationanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Over 698 million references in Crossref lack DOIs, severely impeding the structural construction of citation networks. This paper proposes a systematic approach integrating heuristic rules, metadata matching, and fuzzy text parsing to accurately align unstructured references to target文献 entities in OpenCitations Meta. Methodologically, it combines an interpretable rule engine with semantic matching strategies—specifically designed to enhance linkage accuracy under sparse or inconsistent metadata conditions. Contributions include: (1) a novel hybrid alignment framework balancing precision and robustness; (2) a manually curated gold standard and a validated Crossref subset for rigorous evaluation. Experimental results demonstrate high precision, substantially expanding both the coverage breadth and linkage quality of the open citation network. The method provides a reproducible, scalable technical pathway for large-scale reference DOI enrichment.

Addressing metadata inconsistencies in Crossref references using heuristic toolsCreating formal citation links when references lack specified DOIsMatching bibliographic references with incomplete metadata to existing entities

This work addresses persistent issues in academic manuscripts—such as erroneous citation identifiers, missing metadata, misattributed authorship, and confusion between preprints and published versions—exacerbated by the propensity of large language models to generate citation hallucinations. To mitigate these challenges, we propose a TypeScript-based Model Context Protocol (MCP) server that integrates automated citation validation into intelligent scholarly editing workflows for the first time. Our system employs a manifestation-aware matching mechanism and policy-gated rewriting strategies, harmonizing data from multiple sources including PubMed, Crossref, arXiv, and Semantic Scholar. It supports structured parsing of diverse file formats and multi-round retrieval to generate precise correction suggestions. The prototype has been rigorously evaluated across 47 test cases covering repair actions, exception handling, and protocol compliance, demonstrating robust defense against both conventional citation errors and LLM-induced hallucinations.

bibliographic errorscitation hallucinationLLM-induced errors

Existing benchmarks for reference extraction and parsing are largely confined to well-formatted English bibliographies, struggling to handle the multilingualism, footnote-embedded citations, abbreviations, and diverse historical citation styles prevalent in social sciences and humanities. This work introduces the first unified evaluation benchmark tailored to this domain, comprising three real-world datasets. It systematically assesses the performance of large language models—including DeepSeek-V3.1, Mistral-Small, Gemma-3, and Qwen3-VL—on reference extraction, parsing, and end-to-end document processing tasks, with GROBID serving as a strong baseline. The study further proposes a hybrid deployment strategy combining lightweight LoRA fine-tuning with task routing, which achieves near-saturated extraction performance on medium-sized models and significantly enhances parsing accuracy and system robustness in complex citation scenarios.

footnote citationsmultilingual citationsreference extraction

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

Apr 08, 2024
SS
Sowmya S. Sundaram
🏛️ Stanford University

This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.

Enhance metadata standards adherenceImprove metadata curation automationIntegrate structured knowledge with LLMs

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Latest Papers

What's happening recently
View more

This study addresses a systemic flaw in the citation structures of current retrieval-augmented large language models, which often produce “confirmatory misinformation”—citations that are factually accurate yet structurally distorted, thereby misleading users without their awareness. To tackle this issue, the authors introduce the CITETRACE dataset and propose the first generalizable three-dimensional evaluation framework that comprehensively assesses citation quality across intent alignment, source appropriateness, and answer faithfulness. Leveraging large-scale real-world queries, expert rating matrices, a five-point faithfulness scale, and cross-model citation chain tracing, the analysis reveals that 30.6% of citations distort original sources, 27.1% stem from domain-mismatched references, and 96% of users encounter at least one instance of structural misrepresentation. Citation quality is found to be predominantly influenced by differences among model providers.

citation failurecitation qualitysearch-augmented LLMs

This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.

academic data trackingdata citationdataset usage monitoring

Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.

academic writinghallucinated citationsreference reliability

Existing citation extraction tools struggle to process footnotes in humanities and legal scholarship due to their embedded placement within main text, inclusion of commentary and cross-references, and highly variable formatting. To address this challenge, this work introduces FOSSIL, the first multilingual open dataset specifically designed for footnote citations, comprising over 7,600 annotated references from 96 scholarly papers. The authors also develop a dedicated PDF-TEI Editor collaborative annotation platform, standardize a seven-annotator labeling protocol, and implement a Grobid-based customized footnote parsing model. This end-to-end pipeline substantially improves performance, increasing the micro F1-score from 0.36 to 0.72 with notable gains in recall, thereby demonstrating the approach’s effectiveness while highlighting remaining challenges in handling cross-references and mixed-content footnotes.

bibliographic datacitation extractionfootnotes

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
DK

Daniel Khashabi

Johns Hopkins University
Natural Language ProcessingArtificial IntelligenceMachine Learning
SK

Sadamori Kojaku

Binghamton University
Network scienceComputer scienceScience of Science
MT

Mike Thelwall

School of Information, Journalism and Communication, The University of Sheffield
scientometricsaltmetricssentiment analysissocial media
JB

Joy Bose

Senior Data Scientist at Ericsson, previously in Samsung, Microsoft, Embibe
Machine learningspiking neural networksEEG/BCIlarge language models