citation generation

Designs and builds systems and pipelines that retrieve, select, verify, ground, format, and manage authoritative source references used to support generated text; work includes retrieval of candidate documents and exact sections/subsections, aligning generated claims to specific source passages (citation grounding), removing or preventing hallucinated citations, and organizing citation sets and metadata to maximize recall and precision.

citationgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$179K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.

academic writinghallucinated citationsreference reliability

This work addresses the pervasive issue of citation hallucinations in scientific text generated by large language models, which often manifest as metadata errors or entirely fabricated references. To tackle this challenge, the authors propose CiteCheck, a novel framework that integrates academic database retrieval, a structured large language model verifier, and calibrated multi-tier decision rules to enable fine-grained detection of citation inaccuracies—ranging from minor deviations to complete fabrications. Evaluated on a benchmark comprising 982 physics citations, CiteCheck achieves a macro F1 score of 88.7 and an accuracy of 88.9%, substantially outperforming mainstream models such as GPT, Claude, and Gemini.

citation hallucinationlarge language modelsmetadata corruption

This work addresses the critical challenge of hallucinated citations in scientific texts generated by large language models, which pose a serious threat to academic integrity and are difficult to detect manually. To this end, the authors introduce the first benchmark and verification framework specifically designed for detecting fabricated references in scientific writing. The framework features a novel unified metric for evaluating citation faithfulness and evidence alignment, supported by a large-scale, cross-domain dataset rigorously validated by human annotators. Methodologically, it proposes an interpretable and scalable multi-agent verification pipeline that integrates claim extraction, evidence retrieval, passage matching, and reasoning calibration to enable end-to-end auditing of citation authenticity. Experimental results demonstrate that the proposed approach significantly outperforms existing methods in both accuracy and interpretability, effectively identifying citation errors produced by state-of-the-art large language models and offering a reliable tool for scientific publishing integrity.

citation verificationhallucinated citationsLLM

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Existing citation extraction tools struggle to process footnotes in humanities and legal scholarship due to their embedded placement within main text, inclusion of commentary and cross-references, and highly variable formatting. To address this challenge, this work introduces FOSSIL, the first multilingual open dataset specifically designed for footnote citations, comprising over 7,600 annotated references from 96 scholarly papers. The authors also develop a dedicated PDF-TEI Editor collaborative annotation platform, standardize a seven-annotator labeling protocol, and implement a Grobid-based customized footnote parsing model. This end-to-end pipeline substantially improves performance, increasing the micro F1-score from 0.36 to 0.72 with notable gains in recall, thereby demonstrating the approach’s effectiveness while highlighting remaining challenges in handling cross-references and mixed-content footnotes.

bibliographic datacitation extractionfootnotes

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing evaluation methods for related work generation, which predominantly rely on text similarity metrics and fail to capture critical issues in scholarly positioning, such as inappropriate citations or misaligned references. The paper reframes related work generation as an academic positioning task and introduces RWGBench, a citation-centered, multi-dimensional evaluation framework encompassing four key dimensions: citation selection, contextual appropriateness, organizational structure, and discursive logic. Constructed from 40,108 computer science papers and over 1 million supporting documents, the benchmark combines automated metrics with expert assessments to validate its efficacy. Experiments demonstrate that RWGBench reveals systematic citation-level shortcomings in current systems and that its novel metrics align more closely with expert judgments than conventional text similarity measures, thereby offering a more academically grounded evaluation standard for related work generation.

Citation EvaluationLLM EvaluationRelated Work Generation

This work addresses the challenge of automatically generating fact-checking articles grounded in verifiable citations by leveraging claims, veracity labels, and supporting evidence documents. To this end, the authors propose a multi-agent collaborative pipeline that integrates dense retrieval, source-balanced evidence selection, structured content planning, and citation-aware generation. The framework innovatively incorporates a gated self-evaluation mechanism and a natural language inference (NLI)-driven citation auditing module to repair missing citations and automatically eliminate redundant or unsupported references. Experimental results demonstrate that the proposed approach significantly improves citation accuracy and source credibility in the generated articles, thereby validating the effectiveness of jointly optimizing evidence selection, structured generation, and post-hoc citation verification.

citation auditingevidence groundingfact-checking article generation

This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.

academic data trackingdata citationdataset usage monitoring

Hot Scholars

JC

Joseph Chee Chang

Allen Institute for AI (Ai2)
Human-AI InteractionSensemakingIntelligent User InterfacesResearch Support Tools
ZD

Zhicheng Dou

Renmin University of China
Information RetrievalRetrieval Augmented GenerationLarge Language ModelsGenerative IR
MF

Michael Färber

TU Dresden & ScaDS.AI
Natural Language ProcessingMachine LearningKnowledge Graphs
PS

Pao Siangliulue

Allen Institute for AI
Human Computer InteractionArtificial Intelligence
BD

Bhavana Dalvi Mishra

Lead Research Scientist, Allen Institute for Artificial Intelligence
AgentsInteractive reasoningExplainable machine reasoning