Score
Designs and builds systems that automatically generate, refine, prioritize, and provide auditable provenance for candidate hypotheses or conjectures derived from data, models, or prior theory. These systems integrate foundation-model (LM) outputs with grounding into abstract states, symbolic and numeric checks, exact-verification or proof pipelines, uncertainty/confidence estimates, and agentic critique or iteration loops to search, validate, and formalize conjecture spaces.
Scientific workflows increasingly rely on the edge–cloud–HPC continuum, generating large-scale, structurally complex provenance data; existing analysis approaches—based on scripts, SQL, or static dashboards—suffer from poor interactivity and weak semantic understanding. To address this, we propose the first LLM-based agent system specifically designed for workflow provenance analysis. Our method introduces a modular reference architecture and a dedicated evaluation framework, integrating prompt tuning, retrieval-augmented generation (RAG), and natural language-to-structured query translation to enable deep semantic parsing of provenance metadata and generate insights beyond raw log analysis. The system adopts a lightweight, metadata-driven design and supports multiple foundation models—including LLaMA, GPT, Gemini, and Claude. Evaluated on real-world chemical workflows, it achieves significantly higher query accuracy and analytical depth, enabling dynamic, natural language–driven, interactive provenance exploration.
This work addresses the susceptibility of large language models to hallucinations in empirical reasoning—outputs lacking verifiable evidence or formal guarantees. The authors propose EG-VAR, an architecture that uniquely leverages the Lean 4 formal proof kernel as the sole trusted generator of claims. By integrating tool-certified axioms and a source elevation mechanism, EG-VAR ensures every output is bound to a kernel-verified chain of reasoning and tool invocation; otherwise, it abstains and provides a fully traceable audit trail. Evaluated on a TableBench subset, EG-VAR achieves perfect accuracy (120/120), substantially outperforming a 95% baseline. In counterfactual tests, it maintains 100% source fidelity—significantly higher than competing methods (80–90%)—and exhibits remarkably low semantic formalization error rates of 1.7% (Opus) and 3.3% (Sonnet).
Current benchmarks for mathematical reasoning predominantly rely on answer matching, which fails to assess the logical correctness of solution processes. This work proposes a hybrid verification pipeline that integrates automated and interactive validation by leveraging structured prompting to guide large language models in generating verifiable solutions. The framework supports both formal and informal reasoning and interfaces with proof assistants such as Lean 4, enabling even small-scale models (≤8B parameters) to participate effectively in collaborative verification. Through a multi-agent architecture and advanced prompt engineering, the approach substantially reduces false positive rates. Experimental results demonstrate high verification accuracy across multiple datasets, and the codebase along with deployment guidelines has been publicly released.
Current large language model (LLM) agents lack verifiability, debuggability, and auditability, and relying solely on the accuracy of final answers fails to reveal their underlying reasoning. To address this, this work proposes the first unified provenance framework for LLM agents, systematically modeling causal relationships in tool usage, memory access, and environmental interactions. It introduces a comprehensive provenance taxonomy encompassing source, granularity, representation format, and trust functions. By integrating provenance-aware representation modeling, evidence attribution, runtime safeguards, provenance-informed memory management, and trajectory observability analysis, the study shifts the evaluation paradigm from outcome correctness to process accountability. The framework consolidates existing benchmarks to define a clear pathway for process-level trustworthiness assessment and highlights key challenges, including standardized trajectory schemas, semantic-level provenance, and privacy-preserving auditing.
To address the challenge of tracing derivative relationships among large language models (LLMs), this paper introduces the first black-box model provenance verification framework. Unlike prior approaches, it requires no access to model weights or training data—only API-based output queries—and leverages statistical similarity of output distributions to perform model provenance inference. Crucially, it is the first to formalize this task as a multiple hypothesis testing problem, enabling high-confidence detection of derivative relationships. Evaluated on two real-world benchmarks encompassing over 600 models spanning 30M–4B parameters, the framework achieves 90–95% precision and 80–90% recall. Its core contribution is establishing a rigorous black-box provenance paradigm, supporting intellectual property protection, accountability for model misuse, and identification of foundational model issues—thereby providing a deployable technical foundation for LLM governance.
Current large language model (LLM) agents lack an auditable hypothesis-driven reasoning mechanism in scientific discovery, rendering their inference processes opaque and difficult to verify. This work proposes the Hypothesis Evolution Protocol (HEP), which, for the first time, formalizes hypothesis generation, experimental testing, evidence integration, and belief updating—core components of scientific reasoning—as explicit, structured, and auditable operations. Embedded within an LLM agent architecture, HEP enables end-to-end transparent reasoning and supports cross-task generalization. Evaluated on materials science tasks, HEP agents fully implement a closed-loop hypothesis–test–belief cycle and significantly outperform conventional planning-based agents, achieving notable advances in both auditability and generalization capability.
This work addresses the challenge of subtle errors in mathematical reasoning by large language models through a novel multi-agent framework built upon general-purpose code-oriented large language models. The framework employs a coordinator to dynamically orchestrate a customized pipeline for automatically formalizing research-level mathematical theorems in Lean 4. Its key innovation lies in the ability to dynamically extend type definitions and verify auxiliary lemmas without introducing additional axioms. The approach successfully formalizes the core theorems of five STOC papers—two of which rely solely on the Lean kernel—and produces machine-verified proofs for 32 problems on PutnamBench. All formalizations have been expert-reviewed and are publicly released.
This work addresses the lack of auditable, machine-readable provenance for AI contributions in scientific research. It proposes aiprov, a general-purpose AI provenance model extending the PROV-O ontology to encompass the full human-AI collaborative workflow under FAIR principles. By embedding executable AI capabilities, the system automatically logs operations and generates metadata conforming to provenance graph constraints. Integration with ORCID identity resolution and continuous integration pipelines enforces two critical invariants: “no orphaned claims” and “only humans may authorize validation.” The project itself serves as a complete use case, demonstrating end-to-end traceability and self-auditing of AI-assisted scientific processes.