🤖 AI Summary
Large language models are prone to generating hallucinations or dubious citations in academic writing, undermining research credibility. This study presents the first systematic evaluation and comparison of mainstream citation verification tools—CheckIfExist, HalluCiteChecker, Hallucinator, HalRef, and RefChecker—on real-world academic documents. The analysis reveals significant limitations in current approaches, particularly concerning citation extraction accuracy, breadth of database coverage, and consistency in verification. While these tools can offer preliminary alerts for potentially fabricated references, their overall effectiveness remains constrained. This work provides an empirical foundation and clear directions for improving the verification of citation authenticity in scholarly communication.
📝 Abstract
Large language models are increasingly used in academic writing, including for reference generation, raising concerns about hallucinated and unreliable citations. Recent research suggests that this problem is already widespread and is becoming increasingly prevalent in the published literature and at scientific conferences. In this position paper, we review recent studies on hallucinated references and evaluate several currently available tools for detecting problematic references using documents containing hallucinated citations. The tools assessed include CheckIfExist, HalluCiteChecker, Hallucinator, Hallucinatory Reference Finder (HalRef), and RefChecker. While these systems can provide useful early warnings in many cases, their performance is limited by reference extraction errors, incomplete metadata, limited database coverage, and inconsistent verification results. We argue that hallucinated and suspicious references have become a real and growing problem for scientific communication, and that more transparent and multi-source detection systems are still needed.