Score
Designs and implements graph data structures and generation pipelines that represent evidence items, their provenance and dependency relationships, and any inference constraints; this includes producing hash-linked or otherwise integrity-protected evidence graphs and associated audit/transparency logs. Builds tooling to serialize, validate, and query these graphs so reasoning steps are traceable, structural and constraint-based validity can be checked, and evidence can be used reliably in downstream analyses.
This work addresses the challenge of deploying generative AI in high-stakes decision-making, where hallucinated reasoning, unsupported claims, and weak traceability often preclude compliance with certification-grade accountability requirements. To bridge this gap, the authors propose a “compliance-by-construction” architecture that uniquely integrates typed argumentation graphs, retrieval-augmented generation (RAG), formal verification kernels, and W3C PROV-based provenance tracking. This framework ensures that every AI-generated claim is grounded in authoritative evidence and subjected to rigorous inference constraints before being admitted into official decision records. Empirical evaluation demonstrates that the architecture effectively blocks unsubstantiated assertions from entering the decision pipeline and substantially improves the efficiency of constructing compliant, auditable arguments, thereby enabling controlled, verifiable, and accountable use of generative AI in high-assurance settings.
This work addresses the opacity of reasoning in large language model (LLM) agents performing data-intensive analysis, which hinders the verifiability of their conclusions. To resolve this, the authors propose VeriGraph—a traceable neuro-symbolic reasoning framework that constructs an explicit heterogeneous evidence directed acyclic graph (DAG) to unify raw data, variables, computational results, and natural language claims. VeriGraph introduces three evidence expansion primitives—computation, anchoring, and derivation—to enable structural traceability via graph reachability and incorporates claim-level evidence evaluation to quantify semantic support. Experimental results demonstrate that VeriGraph achieves state-of-the-art performance across four benchmarks, with its 8B variant attaining a claim-level anchoring accuracy of 87.61%, substantially enhancing the auditability and reproducibility of model outputs.
This work addresses critical flaws in scientific reasoning graphs generated by large language models—such as syntactic errors, edge-label drift, incorrect root-node orientation, and weak source anchoring—that render them semantically invalid and non-auditable. To resolve these issues, the authors propose PEARL, a novel framework that, for the first time, integrates Peircean logic–driven closed schemata with an evidence anchoring mechanism. Without requiring model fine-tuning, PEARL employs explicit graph instantiation and a local inference repair algorithm to transform noisy reasoning graphs into semantically valid, fully auditable structures while preserving complete audit trails. Evaluated on the ARCHE benchmark across five datasets comprising 350 scientific papers, PEARL elevates strict compliance from 0 to 300 instances and significantly improves the average Reasoning Graph Effectiveness metric (REA) from 0.339 to 0.906.
This work addresses the open problem of automated verification of cross-paradigm query equivalence between graph databases (Cypher) and relational databases (SQL). We propose the first formal-semantics-based unified verification framework. Our core method introduces a database transformer that directionally embeds a syntactically restricted subset of Cypher into SQL, thereby reducing cross-model equivalence checking to standard SQL query equivalence verification—without requiring joint semantic modeling of both data models. Technically, we integrate formal semantic definitions, syntax-driven translation, SMT solving, and automated theorem proving. We implement our approach in Graphiti, an open-source tool that successfully detects subtle semantic inconsistencies in the official Cypher tutorial and multiple research papers. Graphiti supports efficient equivalence checking and counterexample generation, delivering the first rigorous, formally verifiable, and practically applicable solution for automated cross-paradigm query equivalence verification.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
This work presents the first systematic approach to instance-free schema inference under property graph query transformations. Given a ProGS input schema and a G-CORE query, the authors propose a multi-layer mapping technique that translates property graphs, schemas, and queries into RDF, SHACL, and SPARQL CONSTRUCT representations, respectively, enabling automatic derivation of structural constraints on the output graph via description logic reasoning. By leveraging RDF reification and cross-language semantic bridging, the method establishes a sound and semantically equivalent metatheoretical foundation. This enables generic output schema inference applicable to any input graph conforming to the given schema, while formally verifying both the correctness of the derived constraints and the semantic fidelity of the mappings.