Score
Designs, builds, and evaluates models, automatic metrics, and end-to-end pipelines that detect, verify, and quantify factual correctness of system outputs by comparing claims to source material or external knowledge and flagging unsupported or hallucinated statements. Work includes developing and validating factuality/detection models and metrics, aligning scores with human judgments, calibrating decision thresholds, implementing fact‑grounding and verification components, and extending evaluation approaches to multimodal outputs (e.g., visual or video) when applicable.
Large language models (LLMs) are susceptible to factual hallucinations induced by erroneous information in training data, undermining their reliability. Method: This paper systematically surveys factuality evaluation methodologies, addressing three core challenges: hallucination detection, limitations of existing benchmark datasets, and the reliability of evaluation metrics. We formulate five key research questions and propose a domain-customized fact-checking framework integrating instruction tuning, retrieval-augmented generation (RAG), multi-agent reasoning, and external knowledge integration. Enhanced interpretability and output consistency are achieved via advanced prompting strategies and domain-specific fine-tuning. Contribution/Results: Empirical results demonstrate that evidence-aligned evaluation—leveraging external verifiable sources—significantly outperforms purely autoregressive metrics in hallucination mitigation. The proposed framework advances the development of high-fidelity, context-aware, and domain-adapted trustworthy language models.
This study addresses critical limitations in existing fact-checking datasets—namely, insufficient multilingual coverage, inadequate multimodal evidence integration, lack of structured annotations, and coarse-grained claim-evidence alignment—which hinder research on interpretable and cross-lingual misinformation detection. To overcome these challenges, the authors propose an end-to-end pipeline that aggregates ClaimReview sources and retrieves full debunking articles, standardizes heterogeneous verdicts, and fuses structured metadata with aligned visual content to construct the first French and German multimodal fact-checking datasets. The work introduces a novel fine-grained evidence categorization scheme coupled with a verdict linkage mechanism, leveraging large language models and multimodal large models to automatically extract evidence and generate explanatory rationales. Evaluations via G-Eval and human assessment confirm that the dataset effectively supports the development of interpretable, evidence-based fact-checking models, establishing a foundation for multilingual, multimodal misinformation research.
本文通过引入一种元评估框架,使用控制下的答案扰动测试事实性评估方法的敏感性和可靠性,揭示了基于流水线的方法比LLM作为评判者的方法更能准确追踪退化。
Large language models (LLMs) frequently generate hallucinated content and lack verifiable citations when producing fact-intensive text. To address this, we propose a multi-stage self-verification framework that orchestrates a sequential pipeline of *fact verification → reflective revision → citation integration*. The method synergistically combines chain-of-thought (CoT) reasoning with dual knowledge validation—leveraging both internal consistency checks and external authoritative source alignment—to dynamically perform fine-grained factual scrutiny during generation. When inconsistencies are detected, the model triggers reflective revision and automatically annotates traceable, context-aligned citations. Compared to state-of-the-art approaches, our framework substantially reduces hallucination rates while improving factual accuracy and citation reliability. Empirical evaluation demonstrates its effectiveness in high-fidelity applications such as scientific writing and news generation, where trustworthiness and evidential grounding are critical.
Existing AI-generated content (AIGC) fact-checking tools predominantly rely on black-box binary classification or regression models, suffering from poor interpretability, limited evidence diversity, and minimal user interactivity. Method: We propose the first user-driven, fine-grained fact verification framework that decomposes long texts into atomic claims, integrates heterogeneous multi-source evidence (e.g., knowledge bases, web pages, documents), and employs cross-source evidence fusion with an interpretable reasoning model to produce claim-level confidence scores and natural-language explanations—supporting multi-hop provenance tracing and dynamic user feedback. Contribution/Results: Our framework breaks from conventional paradigms by enabling transparent, traceable, evidence-diverse, and human-AI collaborative verification. Experiments demonstrate significant improvements in user verification efficiency (+37%) and trust (+42%), establishing a novel paradigm for trustworthy AIGC interaction.
Current LLM factuality evaluation suffers from the absence of standardized benchmarks and comparable methodologies, hindering systematic progress. To address this, we propose OpenFactCheck—the first open-source, scalable, and reproducible end-to-end fact-checking framework. Methodologically, it establishes an integrated ecosystem comprising: (i) customizable checker development (CUSTCHECKER), (ii) a cross-model fair evaluation protocol (LLMEVAL), and (iii) human-annotated quantification of checker reliability (CHECKEREVAL); introduces multi-granularity metrics and a human-in-the-loop verification paradigm; and releases standardized benchmark datasets and tooling. Empirically, OpenFactCheck significantly improves the verifiability of LLM outputs, enhances comparability and reliability across diverse fact-checking systems, and provides a unified infrastructure for factuality assessment of open-domain free-text claims.
论文探讨了自动事实核查系统缺乏可验证证据的问题,提出使用形式化方法来解决,并从五个层面组织和分析现有研究,指出当前的不足及未来的研究方向。
Current AI evaluation frameworks overemphasize output correctness while neglecting the resource costs required to verify errors in real-world deployment, allowing high accuracy metrics to mask substantial verification burdens. This work introduces verification-cost errors (VCEs)—errors that cannot be detected by a specified proportion of validators within a given verification budget—thereby shifting the paradigm from defining errors solely by output properties to centering on their detectability during verification. Through an operational definition, verification budget modeling, and user studies, we empirically demonstrate in code generation and multimodal document understanding tasks that high benchmark accuracy can coexist with significant verification effort, underscoring that correctness alone is insufficient to reflect system reliability in practical settings.
This work addresses critical limitations in existing automatic fact-checking benchmarks, which are often constrained in scope, modality, language coverage, and types of misinformation, and further compromised by static datasets vulnerable to data leakage from large model pretraining. To overcome these challenges, we propose VeriTaS—the first dynamic, multimodal fact-checking benchmark—built via a seven-stage automated pipeline that continuously ingests real-world claims from 108 global fact-checking organizations. VeriTaS supports both textual and audiovisual content, incorporates a standardized and decoupled expert adjudication mapping mechanism, and features a fully automated quarterly update strategy. The benchmark encompasses 24,000 claims across 54 languages, with human evaluations confirming high alignment between automated annotations and human judgments, thereby offering a robust, leakage-resistant, and sustainable evaluation platform for the era of large language models.
This work addresses the critical gap in existing text-to-image generation models, which often lack effective mechanisms to ensure factual correctness—particularly for scientific, historical, product-related, or cultural content—and struggle with facts that are implicit in prompts or require external knowledge. To tackle this, the authors propose FAGER, a novel framework that leverages large language models to extract structured facts and generate corresponding question-answer pairs. FAGER integrates reference image–guided visual verification with vision-language models to enable fine-grained, interpretable assessment of factual consistency. Notably, it establishes the first training-free closed-loop generation-feedback optimization pipeline for enhancing factuality. Experiments across five benchmark datasets demonstrate that FAGER significantly outperforms existing metrics in factual A/B testing and effectively improves the factual accuracy of generated images.
Current fact-checking evaluations are largely confined to claim verification, neglecting critical upstream components such as claim extraction and evidence retrieval, thereby failing to comprehensively assess large language models’ systematic reasoning capabilities and factual robustness. To address this limitation, this work proposes FactArena—the first competitive evaluation framework encompassing the entire fact-checking pipeline. FactArena enables automated assessment across all stages—claim decomposition, evidence acquisition, and judgment reasoning—through LLM-driven process standardization, tool-augmented evidence retrieval, a multi-agent adjudication consensus mechanism, and semantically controlled adversarial claim generation. Experiments on 16 mainstream large language models reveal a significant gap between end-to-end fact-checking performance and static verification accuracy, underscoring the necessity and efficacy of holistic, full-pipeline evaluation.