track data lineage

Designs and builds systems, models, and tooling to capture, store, map, and query provenance and data-lineage at both per-sample and dataset levels, including audit trails, signed/locked metadata, source snapshots, refresh history, and explicit links from outputs to executed code and inputs. Analyzes lineage graphs and provenance records to support auditable queries and scoped exports, preserve regulatory facts, gate completion on validation checks, and reconcile provenance across integrations.

trackdatalineage

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

LINEAGEX: A Column Lineage Extraction System for SQL

May 29, 2025
SH
Shi Heng Zhang
🏛️ Simon Fraser University

Addressing the challenge of automatic column-level data lineage extraction in enterprise data governance—where existing approaches either incur runtime overhead or suffer from insufficient accuracy—this paper proposes LINEAGEX, a lightweight static SQL parser based on intelligent traversal of SQL Abstract Syntax Trees (ASTs). LINEAGEX avoids query execution entirely, achieving high coverage and precision through ambiguity-aware column reference resolution and cross-statement contextual tracking. Implemented in Python, the system integrates an interactive lineage graph visualization frontend and is open-sourced on GitHub. Evaluated on real-world enterprise SQL workloads, LINEAGEX achieves state-of-the-art column-level lineage accuracy, significantly outperforming mainstream tools. Its robust, execution-free analysis effectively supports critical governance tasks including data quality diagnostics, storage optimization, and workflow migration.

Extracts column-level lineage from SQL queriesImproves accuracy and coverage in lineage extractionVisualizes lineage via an interactive interface

Learning Lineage Constraints for Data Science Operations

Jun 22, 2025
JZ
Jinjin Zhao
🏛️ University of Chicago

In cross-library data science workflows, existing data lineage representations are tightly coupled to library-specific data models and operational paradigms, hindering tasks such as cross-library debugging. This paper proposes XProv, a unified intermediate representation architecture that integrates concrete data transformation graphs with abstract logical schemas to enable parameterized lineage modeling and correlation analysis for both known and unknown cross-library operations. Inspired by compiler intermediate representations, XProv decouples lineage constraints into extensible logical schemas and supports automatic schema inference from execution traces. Experiments demonstrate that XProv is the first approach to achieve unified logical lineage representation and schema learning across multiple libraries, providing a scalable foundational framework for cross-library provenance tracking, debugging, and verification.

Inferring logical patterns from materialized lineage graphsLinking materialized lineage graphs with abstract logical patternsRepresenting data lineage across diverse libraries uniformly

In-Memory Indexing and Querying of Provenance in Data Preparation Pipelines

Nov 05, 2025
KB
Khalid Belhajjame
🏛️ LAMSADE | Univ. Paris-Dauphine | PSL | Univ. Paris-Dauphine – Tunis

To address the challenges of fine-grained provenance capture and efficient querying in data preparation workflows, this paper proposes a tensor-based in-memory indexing mechanism. The method innovatively integrates both backward- and forward-looking provenance information, employing an enhanced tensor model to explicitly encode record-level and attribute-level input–output mappings, thereby enabling expressive lineage analysis. Unlike conventional approaches, its memory-efficient index design substantially reduces storage overhead while accelerating diverse provenance queries—including origin tracing, impact analysis, and dependency exploration. Experimental evaluation on real-world and synthetic datasets demonstrates that our approach consistently outperforms state-of-the-art baselines in both query latency and memory footprint. The proposed solution thus provides scalable, high-performance provenance support for critical downstream tasks such as debugging, model interpretability, fairness auditing, and data quality diagnostics.

Combining retrospective and prospective provenance for queriesEfficiently capturing and querying data pipeline provenanceMinimizing memory usage while tracking fine-grained lineage

Semantic drift in enterprise data pipelines—caused by multilingual transformations—decouples metadata from downstream data semantics, undermining reproducibility, governance, and performance of RAG and text-to-SQL applications. To address this, we propose a fine-grained schema lineage extraction method leveraging multilingual parsing, chain-of-thought prompting (optimized for 1.3B–32B models), and human-in-the-loop evaluation. We introduce SLiCE (Schema Lineage Composite Evaluation), the first benchmark framework tailored for multilingual script lineage, alongside a high-quality dataset of 1,700 real-world annotated samples. Experiments show that open-weight 32B models match GPT-4’s lineage accuracy under standard prompting, demonstrating cost-effective lineage extraction. Our core contributions are: (1) a systematic formalization of semantic-faithful lineage; (2) the first open, multilingual schema lineage benchmark with rigorous annotations; and (3) a lightweight, efficient extraction paradigm enabling scalable, accurate lineage inference.

Addressing semantic drift in data reproducibility and governanceEvaluating lineage quality with structural and semantic metricsExtracting fine-grained schema lineage from multilingual enterprise pipelines

Existing Model Cards and Data Cards describe only static models and datasets, lacking documentation of the execution context surrounding generation, transformation, and evaluation processes—thereby limiting reproducibility and bias analysis. This work proposes Workflow Cards, which extend the structured documentation paradigm to dynamic workflow executions for the first time. Built upon provenance data, Workflow Cards generate machine-readable, structured summaries interpretable by both humans and large language models (LLMs), and incorporate a template designed to answer typical execution-related questions. Experimental results demonstrate that Workflow Cards substantially enhance understanding of workflow executions compared to schema-based query interfaces, nearly doubling answer quality and achieving superior performance under both LLM-as-a-Judge and human evaluations.

Data CardsModel Cardsprovenance data

Latest Papers

What's happening recently
View more

This study addresses the prevalent issue in computer vision datasets wherein image provenance—such as acquisition parameters and preprocessing steps—is often missing or stored separately, leading to compromised traceability, regulatory compliance, and data reusability. To overcome this limitation, the work proposes a novel approach that leverages JSON-LD (JavaScript Object Notation for Linked Data) to define a structured provenance schema and embeds it directly within image files. This integration ensures an inseparable binding between metadata and the image payload, preserving provenance integrity and persistence across workflows. The proposed method maintains compatibility with existing standards while significantly enhancing data maintainability, system interoperability, and the trustworthiness of downstream models trained on such annotated datasets.

computer vision datasetsdata traceabilityimage provenance

Existing approaches struggle to effectively audit the provenance of skill reuse by large language model agents in multimodal, fragmented scenarios, as evidence is dispersed across textual, code, and operational structures. This work proposes SkillTrace, a novel framework that formulates skill reuse auditing as a multi-trajectory provenance problem. SkillTrace constructs a Skill Operation Graph (SOG) by extracting three types of trajectories—expressive, implementational, and operational—and leverages large language models solely during ingestion for efficient, deterministic trajectory matching. The method incorporates a negative-sample calibration mechanism, achieving an AUROC of 0.938 and an F1 score of 0.898 on SkillTrace-Bench. Evaluation on 36,446 real-world skills reveals actionable instances of skill reuse that surpass repository-level baselines.

LLM-agentmulti-traceoperational graph

Current large language model (LLM) agents lack verifiability, debuggability, and auditability, and relying solely on the accuracy of final answers fails to reveal their underlying reasoning. To address this, this work proposes the first unified provenance framework for LLM agents, systematically modeling causal relationships in tool usage, memory access, and environmental interactions. It introduces a comprehensive provenance taxonomy encompassing source, granularity, representation format, and trust functions. By integrating provenance-aware representation modeling, evidence attribution, runtime safeguards, provenance-informed memory management, and trajectory observability analysis, the study shifts the evaluation paradigm from outcome correctness to process accountability. The framework consolidates existing benchmarks to define a clear pathway for process-level trustworthiness assessment and highlights key challenges, including standardized trajectory schemas, semantic-level provenance, and privacy-preserving auditing.

auditabilityevidence tracingexecution provenance

This work addresses the security risks in multi-agent workflows arising from uncontrolled access to shared memory due to insufficient provenance-aware permissioning and risk awareness. The authors propose a typed execution graph framework that unifies the modeling of agents, provenance, memory, claims, and actions, and—novelty—treats provenance as a runtime control signal to explicitly distinguish hard authorization from graded trust. By integrating lineage tracing, permission filtering, semantic retrieval, multiplicative path-based trust scoring, and a risk-sensitive gating mechanism, the system dynamically adapts evidence usage policies under high-risk actions. Evaluated on 2,700 synthetic tasks, the approach achieves a 94.96% task success rate, 72.70% decision accuracy, and 90.22% clean success rate. Ablation and transfer experiments further validate the contribution of each component.

authorizationmulti-agent workflowsprovenance

Hot Scholars

SP

Silvio Peroni

University of Bologna
Semantic PublishingSemantic WebOpen ScienceScience of Science
DS

Dingjie Song

Lehigh University; CUHK-Shenzhen; Nanjing University
Multimodal LearningLarge Language Models
RF

Rafael Ferreira da Silva

Oak Ridge National Laboratory
Scientific WorkflowsDistributed ComputingWorkflow ManagementModeling and Simulation
CH

Conghui He

Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence