Score
Design, build, or analyze systems that aggregate and synthesize information extracted from multiple documents into structured, claim-level representations (claim graphs / evidence graphs), producing discrete claim variable values and confidence scores. This includes clustering consistent assertions, resolving contradictions and entailments between claims, linking multi-source support, and formatting outputs for downstream consumption such as models or databases.
Existing fact-checking methods suffer from insufficient statement decomposition and ambiguous coreference resolution in LLM-driven statement parsing, leading to high verification complexity and low accuracy. To address these issues, this paper introduces the first approach that models statements as subject–predicate–object (SPO) triple-based graph structures, enabling fine-grained decomposition and explicit coreference disambiguation through graph construction. We propose a relation-constrained, graph-guided reasoning planning mechanism that supports interpretable, triple-wise verification. Furthermore, we design a triple-level semantic alignment framework integrated with LLM-coordinated graph operations. Evaluated on three benchmark datasets—FEVER, SciFact, and Climate-FEVER—our method achieves state-of-the-art performance, significantly improving fine-grained verification accuracy and robustness.
This work addresses the limitations of existing fact-checking approaches, which are often confined to unimodal text or lack interpretability, thereby struggling to verify claims requiring joint visual and textual evidence. The authors propose a two-tier multimodal graph architecture that enables fine-grained evidence retrieval through a bidirectional image-text reasoning mechanism. Multimodal information is fused at both token and evidence levels to support claim verification, while a dedicated fusion decoder generates natural language explanations—realizing, for the first time, an integrated pipeline for retrieval, verification, and explanation. Key contributions include a novel multi-granularity fusion strategy, the bidirectional reasoning mechanism, and AIChartClaim, the first multimodal claim dataset centered on scientific figures in the AI domain. Experiments demonstrate that the proposed method significantly outperforms current baselines on multimodal claim verification tasks.
This work proposes a novel approach to open-domain claim verification that addresses the limitations of existing fact-checking systems, which often rely on a single knowledge source and thus fail to capture perspective divergence, resulting in limited coverage and poor transparency. The method leverages large language models (LLMs) to simultaneously retrieve multi-source evidence—such as from Wikipedia, PubMed, and Google—for both the original claim and its negation, integrating supporting and contradicting information. It further incorporates cross-source disagreement analysis to better model the complexity and diversity of information. By combining evidence deduplication, confidence scoring, and visualization, the approach is evaluated across four benchmark datasets using five distinct LLMs, achieving significant accuracy improvements and revealing notable differences in how various knowledge sources contribute to reasoning.
This study addresses the instability of existing claim decomposition approaches in verifying complex, multifaceted claims, which stems primarily from insufficient evidence alignment and inadequate modeling of error patterns in sub-claims. The authors introduce a new dataset featuring temporally constrained evidence and human-annotated evidence spans for sub-claims, enabling a systematic evaluation of claim decomposition under two evidence alignment settings. They reveal, for the first time, the critical impact of fine-grained evidence alignment and label bias in sub-models on verification performance, and propose a structured decomposition framework that contrasts sub-claim aligned evidence (SAE) with repeated claim-level evidence (SRE). Experiments demonstrate that decomposition significantly improves performance only under strictly aligned, fine-grained evidence conditions, and that incorporating an “abstain” strategy effectively mitigates error propagation, with consistent results across multiple datasets.
This work investigates large language models’ (LLMs) capabilities in evidence-based claim verification, specifically evaluating deductive versus abductive reasoning. To this end, we introduce RECV—the first benchmark featuring real-world claims with fine-grained, atomic-level annotations of reasoning types—and propose a reasoning-type decomposition evaluation framework. Through systematic assessment of mainstream closed-source LLMs across multiple difficulty levels and prompting strategies, complemented by semantic similarity analysis, we find that: (1) LLMs exhibit robust performance on deductive reasoning but suffer from systematic failures in abductive reasoning; (2) generated explanations achieve high semantic similarity to human-written ones—especially for deductive tasks—but rationalization does not consistently improve verification accuracy. This study provides the first empirical evidence of LLMs’ fundamental limitations in abductive reasoning, establishing a novel, trustworthy benchmark and methodology for rigorous reasoning evaluation.
This work proposes a graph neural network–based approach to support compliance and safety verification by analyzing the structural soundness and source authenticity of assurance cases. For the first time, assurance cases are systematically modeled as textual attributed graphs, enabling a joint learning framework that performs link prediction—assessing structural coherence—and human–machine generated case classification—discriminating origin authenticity. The method reveals significant differences in hierarchical linkage patterns between large language model–generated cases and those authored by humans. Experimental results on real-world datasets demonstrate the effectiveness and novelty of the proposed framework, achieving a ROC-AUC of 0.760 for link prediction and an F1 score of 0.94 for human–machine case classification.
Existing approaches struggle to capture fine-grained interactions among scientific papers centered on specific claims. This work proposes ClaimFlow, the first claim-centric NLP framework for scholarly document analysis. By manually annotating 1,084 claims and 832 cross-paper relationships across 304 papers from the ACL Anthology, we define and construct a novel claim relationship classification task aimed at inferring the scientific stance of citing papers toward cited claims. Combining neural architectures with large language models, our approach achieves a macro F1 score of 0.78 on this task. Extending the analysis to 13,000 papers reveals that 63.5% of claims are never reused, only 11.1% are challenged, and widely disseminated claims are typically reframed through qualification or extension.
This work addresses the limitations of existing unsupervised multimodal entity linking methods, which overly rely on instance-centric features while neglecting multi-perspective evidence and their intricate interdependencies. To overcome this, the authors propose MSR-MEL, a novel framework that integrates offline multi-view evidence synthesis with online large language model (LLM) reasoning. MSR-MEL innovatively introduces group-level evidence, leverages contextualized graph structures to aggregate neighborhood information, and combines an asymmetric teacher–student graph neural network with LLM-based semantic reasoning. Evaluated on standard benchmarks, MSR-MEL substantially outperforms current unsupervised approaches, achieving significantly higher entity linking accuracy.
This work addresses the opacity of reasoning in large language model (LLM) agents performing data-intensive analysis, which hinders the verifiability of their conclusions. To resolve this, the authors propose VeriGraph—a traceable neuro-symbolic reasoning framework that constructs an explicit heterogeneous evidence directed acyclic graph (DAG) to unify raw data, variables, computational results, and natural language claims. VeriGraph introduces three evidence expansion primitives—computation, anchoring, and derivation—to enable structural traceability via graph reachability and incorporates claim-level evidence evaluation to quantify semantic support. Experimental results demonstrate that VeriGraph achieves state-of-the-art performance across four benchmarks, with its 8B variant attaining a claim-level anchoring accuracy of 87.61%, substantially enhancing the auditability and reproducibility of model outputs.