Score
Design, build, or analyze methods that identify and delineate individual (atomic) claims in documents by detecting claim-bearing sentences, segmenting and mapping claim tokens to source spans, and distinguishing central atoms from background text and qualifiers. Produce structured, atom-level representations or verification units by decomposing complex claims into empirical atoms, filtering candidates by in-paper context, and linking claims to specific citation contexts for downstream verification and reasoning.
To address challenges in complex claim verification—including multi-hop reasoning difficulties, error propagation, and evidence noise—this paper proposes a dynamic iterative atomic decomposition framework. First, context-aware atomic facts are extracted using large language models; then, fine-grained, adaptive multi-hop reasoning is achieved through semantic refinement and evidence re-ranking. Crucially, the framework innovatively integrates prompt-guided chain-of-thought reasoning with iterative verification, effectively mitigating structural modeling deficiencies and intent misalignment. Evaluated on five mainstream benchmarks, the method achieves state-of-the-art accuracy while significantly enhancing interpretability and robustness to erroneous or noisy evidence.
Document-level claim extraction for fact-checking lacks robust evaluation methodologies—particularly for low-resource languages (e.g., Czech and Slovak) and informal text. To address this, we propose the first multilingual, semantics-aware evaluation framework for document-level claim extraction, grounded in three core criteria: atomicity, verifiability, and decontextualization. Our framework enables reliable comparison between model outputs and human annotations via claim-set alignment and fine-grained semantic similarity computation. Experiments on a newly curated Czech/Slovak news commentary dataset demonstrate that conventional metrics (e.g., precision/recall) severely underestimate model performance. In contrast, our approach more accurately reflects model capabilities, quantifies inter-annotator agreement, and establishes a reproducible, interpretable evaluation paradigm for cross-lingual fact-checking.
This study addresses the limitations of existing verifiable claim detection methods, which rely solely on claim text and neglect contextual information, thereby constraining verification reliability. To overcome this, the work introduces context retrieval into the task for the first time and proposes an end-to-end, context-driven detection paradigm. Specifically, it retrieves relevant evidence from Wikipedia via entity recognition and leverages large language models to generate contextual summaries that support claim classification. Experiments on the CheckThat! 2022 and PoliClaim datasets demonstrate the effectiveness of the approach, revealing that its performance is influenced by domain characteristics, model architecture, and learning settings—including fine-tuning, zero-shot, and few-shot configurations. Multidimensional analysis further elucidates the mechanisms through which contextual augmentation enhances verification accuracy.
Large language models (LLMs) frequently generate fabricated references in scholarly writing—such as fictitious authors or incorrect DOIs—a problem that has already infiltrated top-tier conferences and journals and often evades detection by conventional peer review. This work proposes the first automated detection tool that integrates LLM-based field extraction with structured querying of academic databases. Specifically, the method employs an LLM to parse citation fields and then leverages Semantic Scholar to perform semantic matching and similarity scoring based on title, author names, and publication venue, yielding a tiered credibility assessment (credible, partially supported, or likely fabricated). Evaluated on a manually annotated test set of papers accepted at NeurIPS 2025, the approach efficiently identifies the vast majority of hallucinated citations, offering a scalable technical safeguard for research integrity.
Scientific claims on social media are often difficult to trace back to their original sources due to variations in language, style, and detail, posing a significant challenge for automated fact-checking. This work systematically evaluates sparse and dense retrieval models on the CheckThat! 2026 benchmark, integrating multilingual translation, publication metadata, four style-transfer strategies, and re-ranking techniques. The study finds that translating claims into English yields better performance than using original or bilingual representations. It introduces three novel re-ranking models based on attribution, entity overlap, and verification-driven reasoning, with the latter significantly outperforming semantic similarity baselines and achieving a state-of-the-art MRR@5 of 0.758. Furthermore, the effectiveness of style-transfer strategies is shown to depend critically on the retrieval objective.
This work addresses the challenge of synthesizing fragmented chemical knowledge from vast scientific literature, a process traditionally reliant on manual curation and poorly supported by conventional document-retrieval systems. To overcome this limitation, the authors propose a novel infrastructure centered on “claims”—structured, traceable assertions—as the fundamental unit of retrieval, shifting from whole-document indexing to fine-grained knowledge representation. The system extracts claims from publications, anchors them to their original sources, classifies them via faceted taxonomies, and organizes them into evidence graphs for dynamic knowledge integration. Deployed over 147,000 papers, it indexes 2.4 million claims, substantially increasing citation density and enabling GPT-5.5 to achieve 100% DOI-resolvability on the AskChem-Bench benchmark. This framework provides an efficient, interpretable interface for both AI agents and researchers to collaboratively explore and validate scientific knowledge.
This work addresses limitations in existing compound claim decomposition methods for automated fact-checking, which rely on lexical overlap metrics like Jaccard similarity and thus struggle to accurately assess the semantic fidelity of paraphrased atomic claims, while also lacking theoretical guarantees on the termination of repair processes. To overcome these issues, the authors propose CREDENCE, a framework that replaces lexical overlap with cosine similarity based on BGE-large embeddings to enable semantic-aware decomposition and self-repair. They formally prove, for the first time, the convergence of a hybrid repair pipeline combining symbolic rules and large language models. The study introduces Semantic-F1, a new evaluation metric validated across three cross-domain benchmarks—social media, encyclopedic texts, and news—demonstrating 15–32 percentage point improvements over Jaccard-F1, EPR scores of 0.94–1.00, and a 47%–100% reduction in atomicity violations via rule-based repair without compromising semantic fidelity.