Score
Design, build, or evaluate systems that generate structured reports whose individual claims are explicitly linked to verifiable evidence and recorded provenance, so every finding is traceable. Work includes mechanisms for claim–evidence linking, graph- or knowledge-graph-grounded generation and in-context grounding, and for detecting or flagging statements that lack supporting evidence.
The evidence-based text generation (EBTG) field suffers from terminological ambiguity, fragmented evaluation practices, and the absence of standardized benchmarks. Method: We conduct a systematic review of 134 papers and propose the first unified taxonomy covering three core operations—citation, attribution, and quotation—structured along seven dimensions (e.g., evidence source, granularity, explicitness). Concurrently, we design a multidimensional evaluation framework that integrates and standardizes 300 metrics, and introduce the first reproducible, comparable EBTG benchmark. Contribution/Results: Our work resolves long-standing fragmentation in EBTG research, establishes a rigorous foundation for assessing output traceability, verifiability, and trustworthiness, and provides both theoretical guidance and practical tools to advance reliable large language model development.
This work addresses the challenge of deploying generative AI in high-stakes decision-making, where hallucinated reasoning, unsupported claims, and weak traceability often preclude compliance with certification-grade accountability requirements. To bridge this gap, the authors propose a “compliance-by-construction” architecture that uniquely integrates typed argumentation graphs, retrieval-augmented generation (RAG), formal verification kernels, and W3C PROV-based provenance tracking. This framework ensures that every AI-generated claim is grounded in authoritative evidence and subjected to rigorous inference constraints before being admitted into official decision records. Empirical evaluation demonstrates that the architecture effectively blocks unsubstantiated assertions from entering the decision pipeline and substantially improves the efficiency of constructing compliant, auditable arguments, thereby enabling controlled, verifiable, and accountable use of generative AI in high-assurance settings.
This work addresses the opacity of reasoning in large language model (LLM) agents performing data-intensive analysis, which hinders the verifiability of their conclusions. To resolve this, the authors propose VeriGraph—a traceable neuro-symbolic reasoning framework that constructs an explicit heterogeneous evidence directed acyclic graph (DAG) to unify raw data, variables, computational results, and natural language claims. VeriGraph introduces three evidence expansion primitives—computation, anchoring, and derivation—to enable structural traceability via graph reachability and incorporates claim-level evidence evaluation to quantify semantic support. Experimental results demonstrate that VeriGraph achieves state-of-the-art performance across four benchmarks, with its 8B variant attaining a claim-level anchoring accuracy of 87.61%, substantially enhancing the auditability and reproducibility of model outputs.
This work investigates large language models’ (LLMs) capabilities in evidence-based claim verification, specifically evaluating deductive versus abductive reasoning. To this end, we introduce RECV—the first benchmark featuring real-world claims with fine-grained, atomic-level annotations of reasoning types—and propose a reasoning-type decomposition evaluation framework. Through systematic assessment of mainstream closed-source LLMs across multiple difficulty levels and prompting strategies, complemented by semantic similarity analysis, we find that: (1) LLMs exhibit robust performance on deductive reasoning but suffer from systematic failures in abductive reasoning; (2) generated explanations achieve high semantic similarity to human-written ones—especially for deductive tasks—but rationalization does not consistently improve verification accuracy. This study provides the first empirical evidence of LLMs’ fundamental limitations in abductive reasoning, establishing a novel, trustworthy benchmark and methodology for rigorous reasoning evaluation.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
This work addresses the frequent failure of autonomous scientific agents due to unsupported claims or inconsistencies across research stages. To tackle this, the authors propose representing the agent’s internal state as a typed evidence graph comprising nodes for questions, knowledge gaps, hypotheses, experiments, findings, and claims. This framework enables, for the first time, real-time tracking of claim–evidence consistency, precise identification of flaws, and targeted repairs, with a graph checkpointing mechanism ensuring the safety of such interventions. The approach integrates evidence graph construction, consistency verification, regeneration of weak nodes, and chain-of-verification–guided paper generation. Evaluated on ARC-Bench-ML and NanoResearch-20, the method improves claim support rate by 40.19% and achieves 87.73% experimental data consistency, substantially outperforming baseline systems.
This work addresses the challenge that AI-generated claims often lack reliable evidential support, while manual verification remains inefficient and poorly traceable. To tackle this, the authors propose an Evidence Ledger Adjudication pipeline—the first to integrate an evidence ledger mechanism into AI-assisted writing—automatically pairing each claim with a heterogeneous evidence bundle and classifying their relationship as supporting, contradicting, or mixed. Problematic claims are then routed back for revision. By constructing a blind-test benchmark based on external annotations and leveraging relation classification alongside evidence provenance techniques, the system enables automated consistency assessment and routing between claims and evidence. Evaluated on 2,335 samples, the approach achieves a relation accuracy of 0.676 and a macro F1 of 0.601, significantly outperforming baselines; it correctly identifies 89.2% of problematic claims while misrouting only 32.8% of valid ones, substantially enhancing the management of claim credibility.