Score
Designs and builds systems, models, and tooling to capture, store, map, and query provenance and data-lineage at both per-sample and dataset levels, including audit trails, signed/locked metadata, source snapshots, refresh history, and explicit links from outputs to executed code and inputs. Analyzes lineage graphs and provenance records to support auditable queries and scoped exports, preserve regulatory facts, gate completion on validation checks, and reconcile provenance across integrations.
Addressing the challenge of automatic column-level data lineage extraction in enterprise data governance—where existing approaches either incur runtime overhead or suffer from insufficient accuracy—this paper proposes LINEAGEX, a lightweight static SQL parser based on intelligent traversal of SQL Abstract Syntax Trees (ASTs). LINEAGEX avoids query execution entirely, achieving high coverage and precision through ambiguity-aware column reference resolution and cross-statement contextual tracking. Implemented in Python, the system integrates an interactive lineage graph visualization frontend and is open-sourced on GitHub. Evaluated on real-world enterprise SQL workloads, LINEAGEX achieves state-of-the-art column-level lineage accuracy, significantly outperforming mainstream tools. Its robust, execution-free analysis effectively supports critical governance tasks including data quality diagnostics, storage optimization, and workflow migration.
In cross-library data science workflows, existing data lineage representations are tightly coupled to library-specific data models and operational paradigms, hindering tasks such as cross-library debugging. This paper proposes XProv, a unified intermediate representation architecture that integrates concrete data transformation graphs with abstract logical schemas to enable parameterized lineage modeling and correlation analysis for both known and unknown cross-library operations. Inspired by compiler intermediate representations, XProv decouples lineage constraints into extensible logical schemas and supports automatic schema inference from execution traces. Experiments demonstrate that XProv is the first approach to achieve unified logical lineage representation and schema learning across multiple libraries, providing a scalable foundational framework for cross-library provenance tracking, debugging, and verification.
To address the challenges of fine-grained provenance capture and efficient querying in data preparation workflows, this paper proposes a tensor-based in-memory indexing mechanism. The method innovatively integrates both backward- and forward-looking provenance information, employing an enhanced tensor model to explicitly encode record-level and attribute-level input–output mappings, thereby enabling expressive lineage analysis. Unlike conventional approaches, its memory-efficient index design substantially reduces storage overhead while accelerating diverse provenance queries—including origin tracing, impact analysis, and dependency exploration. Experimental evaluation on real-world and synthetic datasets demonstrates that our approach consistently outperforms state-of-the-art baselines in both query latency and memory footprint. The proposed solution thus provides scalable, high-performance provenance support for critical downstream tasks such as debugging, model interpretability, fairness auditing, and data quality diagnostics.
Semantic drift in enterprise data pipelines—caused by multilingual transformations—decouples metadata from downstream data semantics, undermining reproducibility, governance, and performance of RAG and text-to-SQL applications. To address this, we propose a fine-grained schema lineage extraction method leveraging multilingual parsing, chain-of-thought prompting (optimized for 1.3B–32B models), and human-in-the-loop evaluation. We introduce SLiCE (Schema Lineage Composite Evaluation), the first benchmark framework tailored for multilingual script lineage, alongside a high-quality dataset of 1,700 real-world annotated samples. Experiments show that open-weight 32B models match GPT-4’s lineage accuracy under standard prompting, demonstrating cost-effective lineage extraction. Our core contributions are: (1) a systematic formalization of semantic-faithful lineage; (2) the first open, multilingual schema lineage benchmark with rigorous annotations; and (3) a lightweight, efficient extraction paradigm enabling scalable, accurate lineage inference.
Existing Model Cards and Data Cards describe only static models and datasets, lacking documentation of the execution context surrounding generation, transformation, and evaluation processes—thereby limiting reproducibility and bias analysis. This work proposes Workflow Cards, which extend the structured documentation paradigm to dynamic workflow executions for the first time. Built upon provenance data, Workflow Cards generate machine-readable, structured summaries interpretable by both humans and large language models (LLMs), and incorporate a template designed to answer typical execution-related questions. Experimental results demonstrate that Workflow Cards substantially enhance understanding of workflow executions compared to schema-based query interfaces, nearly doubling answer quality and achieving superior performance under both LLM-as-a-Judge and human evaluations.
This study addresses the prevalent issue in computer vision datasets wherein image provenance—such as acquisition parameters and preprocessing steps—is often missing or stored separately, leading to compromised traceability, regulatory compliance, and data reusability. To overcome this limitation, the work proposes a novel approach that leverages JSON-LD (JavaScript Object Notation for Linked Data) to define a structured provenance schema and embeds it directly within image files. This integration ensures an inseparable binding between metadata and the image payload, preserving provenance integrity and persistence across workflows. The proposed method maintains compatibility with existing standards while significantly enhancing data maintainability, system interoperability, and the trustworthiness of downstream models trained on such annotated datasets.
Existing approaches struggle to effectively audit the provenance of skill reuse by large language model agents in multimodal, fragmented scenarios, as evidence is dispersed across textual, code, and operational structures. This work proposes SkillTrace, a novel framework that formulates skill reuse auditing as a multi-trajectory provenance problem. SkillTrace constructs a Skill Operation Graph (SOG) by extracting three types of trajectories—expressive, implementational, and operational—and leverages large language models solely during ingestion for efficient, deterministic trajectory matching. The method incorporates a negative-sample calibration mechanism, achieving an AUROC of 0.938 and an F1 score of 0.898 on SkillTrace-Bench. Evaluation on 36,446 real-world skills reveals actionable instances of skill reuse that surpass repository-level baselines.
Current large language model (LLM) agents lack verifiability, debuggability, and auditability, and relying solely on the accuracy of final answers fails to reveal their underlying reasoning. To address this, this work proposes the first unified provenance framework for LLM agents, systematically modeling causal relationships in tool usage, memory access, and environmental interactions. It introduces a comprehensive provenance taxonomy encompassing source, granularity, representation format, and trust functions. By integrating provenance-aware representation modeling, evidence attribution, runtime safeguards, provenance-informed memory management, and trajectory observability analysis, the study shifts the evaluation paradigm from outcome correctness to process accountability. The framework consolidates existing benchmarks to define a clear pathway for process-level trustworthiness assessment and highlights key challenges, including standardized trajectory schemas, semantic-level provenance, and privacy-preserving auditing.
This work addresses the security risks in multi-agent workflows arising from uncontrolled access to shared memory due to insufficient provenance-aware permissioning and risk awareness. The authors propose a typed execution graph framework that unifies the modeling of agents, provenance, memory, claims, and actions, and—novelty—treats provenance as a runtime control signal to explicitly distinguish hard authorization from graded trust. By integrating lineage tracing, permission filtering, semantic retrieval, multiplicative path-based trust scoring, and a risk-sensitive gating mechanism, the system dynamically adapts evidence usage policies under high-risk actions. Evaluated on 2,700 synthetic tasks, the approach achieves a 94.96% task success rate, 72.70% decision accuracy, and 90.22% clean success rate. Ablation and transfer experiments further validate the contribution of each component.