Score
Designs and builds systems and pipelines that retrieve, select, verify, ground, format, and manage authoritative source references used to support generated text; work includes retrieval of candidate documents and exact sections/subsections, aligning generated claims to specific source passages (citation grounding), removing or preventing hallucinated citations, and organizing citation sets and metadata to maximize recall and precision.
The evidence-based text generation (EBTG) field suffers from terminological ambiguity, fragmented evaluation practices, and the absence of standardized benchmarks. Method: We conduct a systematic review of 134 papers and propose the first unified taxonomy covering three core operations—citation, attribution, and quotation—structured along seven dimensions (e.g., evidence source, granularity, explicitness). Concurrently, we design a multidimensional evaluation framework that integrates and standardizes 300 metrics, and introduce the first reproducible, comparable EBTG benchmark. Contribution/Results: Our work resolves long-standing fragmentation in EBTG research, establishes a rigorous foundation for assessing output traceability, verifiability, and trustworthiness, and provides both theoretical guidance and practical tools to advance reliable large language model development.
Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.
This work addresses the pervasive issue of citation hallucinations in scientific text generated by large language models, which often manifest as metadata errors or entirely fabricated references. To tackle this challenge, the authors propose CiteCheck, a novel framework that integrates academic database retrieval, a structured large language model verifier, and calibrated multi-tier decision rules to enable fine-grained detection of citation inaccuracies—ranging from minor deviations to complete fabrications. Evaluated on a benchmark comprising 982 physics citations, CiteCheck achieves a macro F1 score of 88.7 and an accuracy of 88.9%, substantially outperforming mainstream models such as GPT, Claude, and Gemini.
This work addresses the critical challenge of hallucinated citations in scientific texts generated by large language models, which pose a serious threat to academic integrity and are difficult to detect manually. To this end, the authors introduce the first benchmark and verification framework specifically designed for detecting fabricated references in scientific writing. The framework features a novel unified metric for evaluating citation faithfulness and evidence alignment, supported by a large-scale, cross-domain dataset rigorously validated by human annotators. Methodologically, it proposes an interpretable and scalable multi-agent verification pipeline that integrates claim extraction, evidence retrieval, passage matching, and reasoning calibration to enable end-to-end auditing of citation authenticity. Experimental results demonstrate that the proposed approach significantly outperforms existing methods in both accuracy and interpretability, effectively identifying citation errors produced by state-of-the-art large language models and offering a reliable tool for scientific publishing integrity.
This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.
Existing citation extraction tools struggle to process footnotes in humanities and legal scholarship due to their embedded placement within main text, inclusion of commentary and cross-references, and highly variable formatting. To address this challenge, this work introduces FOSSIL, the first multilingual open dataset specifically designed for footnote citations, comprising over 7,600 annotated references from 96 scholarly papers. The authors also develop a dedicated PDF-TEI Editor collaborative annotation platform, standardize a seven-annotator labeling protocol, and implement a Grobid-based customized footnote parsing model. This end-to-end pipeline substantially improves performance, increasing the micro F1-score from 0.36 to 0.72 with notable gains in recall, thereby demonstrating the approach’s effectiveness while highlighting remaining challenges in handling cross-references and mixed-content footnotes.
This work addresses the limitations of existing evaluation methods for related work generation, which predominantly rely on text similarity metrics and fail to capture critical issues in scholarly positioning, such as inappropriate citations or misaligned references. The paper reframes related work generation as an academic positioning task and introduces RWGBench, a citation-centered, multi-dimensional evaluation framework encompassing four key dimensions: citation selection, contextual appropriateness, organizational structure, and discursive logic. Constructed from 40,108 computer science papers and over 1 million supporting documents, the benchmark combines automated metrics with expert assessments to validate its efficacy. Experiments demonstrate that RWGBench reveals systematic citation-level shortcomings in current systems and that its novel metrics align more closely with expert judgments than conventional text similarity measures, thereby offering a more academically grounded evaluation standard for related work generation.
This work addresses the challenge of automatically generating fact-checking articles grounded in verifiable citations by leveraging claims, veracity labels, and supporting evidence documents. To this end, the authors propose a multi-agent collaborative pipeline that integrates dense retrieval, source-balanced evidence selection, structured content planning, and citation-aware generation. The framework innovatively incorporates a gated self-evaluation mechanism and a natural language inference (NLI)-driven citation auditing module to repair missing citations and automatically eliminate redundant or unsupported references. Experimental results demonstrate that the proposed approach significantly improves citation accuracy and source credibility in the generated articles, thereby validating the effectiveness of jointly optimizing evidence selection, structured generation, and post-hoc citation verification.
This work addresses the lack of effective monitoring of dataset usage in scholarly literature, which undermines citation transparency, impact traceability, and reproducibility. To tackle this challenge, the study introduces the first application of the multi-task GLiNER framework to dataset usage monitoring, jointly performing dataset mention extraction, relation identification, and usage context classification. The approach integrates synthetic data generation with a large language model (LLM)-driven re-verification mechanism to mitigate issues of annotation scarcity and ambiguous citations. This combination significantly enhances the accuracy, coverage, and label consistency of dataset mention detection, enabling end-to-end, unconstrained tracking of data citations across diverse scientific texts and advancing the development of open-source tools for scholarly data provenance.