Score
Designs and analyzes systems and tools that record, model, trace, and verify the origin and transformation history of data and metadata—including provenance logs, provenance graphs, dataset and metadata provenance analysis, artifact tracking, and auditing mechanisms. Develops methods to infer and reason about provenance edges and explanations (e.g., evidence-based and LLM-based provenance inference), and to generate proof-grounded, auditable explanations that support provenance verification.
Current large language model (LLM) agents lack verifiability, debuggability, and auditability, and relying solely on the accuracy of final answers fails to reveal their underlying reasoning. To address this, this work proposes the first unified provenance framework for LLM agents, systematically modeling causal relationships in tool usage, memory access, and environmental interactions. It introduces a comprehensive provenance taxonomy encompassing source, granularity, representation format, and trust functions. By integrating provenance-aware representation modeling, evidence attribution, runtime safeguards, provenance-informed memory management, and trajectory observability analysis, the study shifts the evaluation paradigm from outcome correctness to process accountability. The framework consolidates existing benchmarks to define a clear pathway for process-level trustworthiness assessment and highlights key challenges, including standardized trajectory schemas, semantic-level provenance, and privacy-preserving auditing.
This study addresses the lack of transparency and traceability in the training data lifecycle of large language models (LLMs), which undermines their trustworthiness and auditability. Through a systematic review of 95 relevant publications over the past decade, this work proposes the first unified classification framework tailored to the LLM data lifecycle, structured around three core dimensions: data provenance, transparency, and traceability. The framework integrates key technical approaches—including data generation, watermarking, bias measurement, privacy preservation, and governance tools—thereby clarifying the field’s boundaries and revealing inherent trade-offs between transparency and opacity. Furthermore, it synthesizes emerging research trends and open challenges, offering both a theoretical foundation and practical guidance for enhancing the credibility of LLM training data.
This work addresses the challenges of data provenance in streaming systems—namely high computational overhead, substantial storage costs, and limited scalability—which have restricted existing approaches primarily to debugging scenarios and hindered broader data management applications. To overcome these limitations, the study introduces temporal interaction networks (TINs) into the provenance domain for the first time and proposes an efficient time-aware provenance framework tailored for data management. It formally defines two classes of data—discrete and liquid—and five types of temporal provenance queries, accompanied by a state-based indexing mechanism. Experimental evaluations on Apache Flink demonstrate significant reductions in both storage and computational costs, while case studies in transportation and finance underscore the framework’s scalability and practical utility, thereby extending the applicability and performance boundaries of traditional provenance techniques.
Scientific workflows increasingly rely on the edge–cloud–HPC continuum, generating large-scale, structurally complex provenance data; existing analysis approaches—based on scripts, SQL, or static dashboards—suffer from poor interactivity and weak semantic understanding. To address this, we propose the first LLM-based agent system specifically designed for workflow provenance analysis. Our method introduces a modular reference architecture and a dedicated evaluation framework, integrating prompt tuning, retrieval-augmented generation (RAG), and natural language-to-structured query translation to enable deep semantic parsing of provenance metadata and generate insights beyond raw log analysis. The system adopts a lightweight, metadata-driven design and supports multiple foundation models—including LLaMA, GPT, Gemini, and Claude. Evaluated on real-world chemical workflows, it achieves significantly higher query accuracy and analytical depth, enabling dynamic, natural language–driven, interactive provenance exploration.
Digital humanities face challenges in provenance tracking and change management for cultural heritage metadata, as existing RDF-based approaches suffer from weak standards compliance (e.g., W3C RDF reification, n-ary relations) and poor cross-domain interoperability. Method: This study conducts a systematic, multidimensional empirical evaluation of six mainstream semantic models—Named Graphs, RDF*, PROV-O, among others—assessing their standards conformance, extensibility, and domain adaptability specifically within cultural heritage contexts. Contribution/Results: We propose a practice-oriented provenance modeling selection framework that explicitly characterizes trade-offs among trustworthiness assurance, computational overhead, and interoperability. The framework delivers reusable, verifiable decision support for metadata provenance modeling in digital humanities projects, thereby bridging a critical methodological gap in the deep adaptation of Semantic Web technologies to humanities scholarship.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
Existing approaches struggle to effectively audit the provenance of skill reuse by large language model agents in multimodal, fragmented scenarios, as evidence is dispersed across textual, code, and operational structures. This work proposes SkillTrace, a novel framework that formulates skill reuse auditing as a multi-trajectory provenance problem. SkillTrace constructs a Skill Operation Graph (SOG) by extracting three types of trajectories—expressive, implementational, and operational—and leverages large language models solely during ingestion for efficient, deterministic trajectory matching. The method incorporates a negative-sample calibration mechanism, achieving an AUROC of 0.938 and an F1 score of 0.898 on SkillTrace-Bench. Evaluation on 36,446 real-world skills reveals actionable instances of skill reuse that surpass repository-level baselines.
Existing Model Cards and Data Cards describe only static models and datasets, lacking documentation of the execution context surrounding generation, transformation, and evaluation processes—thereby limiting reproducibility and bias analysis. This work proposes Workflow Cards, which extend the structured documentation paradigm to dynamic workflow executions for the first time. Built upon provenance data, Workflow Cards generate machine-readable, structured summaries interpretable by both humans and large language models (LLMs), and incorporate a template designed to answer typical execution-related questions. Experimental results demonstrate that Workflow Cards substantially enhance understanding of workflow executions compared to schema-based query interfaces, nearly doubling answer quality and achieving superior performance under both LLM-as-a-Judge and human evaluations.
Existing automated approaches for mapping cyber threat intelligence (CTI) to MITRE ATT&CK lack supporting evidence, provenance tracking, and validation history, making their credibility difficult to assess. This work proposes the first knowledge graph–driven framework for CTI governance that enables auditable management of TTP assertions through fine-grained evidence preservation, complete provenance chains, versioned trust decisions, and lossless revocation mechanisms. The framework integrates multi-extractor collaborative verification, assertion aggregation, consensus modeling, and policy-driven validation, all underpinned by versioned knowledge graph management. Evaluated on 65 CTI reports comprising 5,303 sentences, the approach achieves a precision of 90.6% under six-party consensus and efficiently supports seven categories of audit queries concerning provenance, trustworthiness, and versioning.