change data capture

Designs, builds, and evaluates systems that detect and extract modifications made to persistent data stores and emit those modifications as ordered change events or logs for downstream consumers. This includes engineering capture mechanisms, schema-evolution handling, delivery semantics (ordering, latency, and at-least/at-most/exactly-once guarantees), and integration with replication, streaming pipelines, or event-sourced materializations.

changedatacapture

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.96
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$230K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Databases continuously evolve through operations such as schema changes, version updates, and data transformations; however, existing approaches typically address these functionalities in isolation, lacking a unified abstraction. This work proposes the first integrated model that unifies continuous schema evolution, version management, and data transformation within a single framework. Built upon general-purpose computational primitives, the model supports operation provenance, conditional update propagation, and change alerts, while employing a declarative mechanism to manage the co-evolution of dependent artifacts—including views and machine learning models. A prototype system implements this framework using an enhanced, parameterized Prolly Tree—a Merkle tree–inspired data structure—to construct a relational-like engine. Experimental evaluation demonstrates that the proposed approach is both feasible and offers tunable performance across diverse evolution scenarios.

continuous data evolutiondata versioningdatabase transformation

Baseline: Operation-Based Evolution and Versioning of Data

Dec 10, 2025
JE
Jonathan Edwards
🏛️ Independent | Charles University

This paper addresses version control challenges for multidimensional structured data—namely, temporal evolution, spatial collaboration, and design iteration—by introducing Operational Differencing, a novel paradigm. Methodologically, it incorporates high-level semantic operations (e.g., schema changes, refactorings) into the version model; adopts an append-only branch history with a repository-free, lightweight “copy-as-branch” architecture; and enables operational query representation and future-tense execution. Contributions include: (1) the first systematic support for automatic schema-adaptive query rewriting under schema evolution; (2) precise, fine-grained diff/merge across structural transformations; (3) resolution of four out of eight canonical schema evolution challenges; and (4) a simplified versioning experience with no explicit repository and asymptotically zero branching overhead.

Automatically adapting queries to accommodate schema changesEnabling version control with fine-grained diffing despite structural transformationsManaging data evolution across time, collaboration, and design changes

This work addresses the problem of providing provably correct snapshot-equivalent change capture and replay for continuously written databases without relying on global snapshots or locking mechanisms. It introduces the “authenticated virtual slice” model, which combines interleaved log scanning with watermark-based positioning to enable incremental authentication over primary key ranges and advancing frontiers while preserving source continuity. For the first time, the paper formally defines and machine-verifies snapshot equivalence for DBLog, presenting a rigorous Isabelle/HOL proof that all well-formed executions satisfy per-key replay equivalence within their designated frontiers and key ranges, and that the authentication process yields valid virtual slices.

change-data-capturecorrectness formalizationdatabase replication

Modern OLTP systems often suffer from frequent schema changes, missing primary/foreign keys, and fragmented execution traces, rendering traditional approaches—reliant on fixed schemas and manual modeling—costly and error-prone. This work proposes a fully automated pipeline that operates without predefined schemas by identifying quasi-key and timestamp columns, discovering inter-table relationships through statistical signals, and assembling and ordering events accordingly. To capture long-range dependencies across system events, the method incorporates a Temporal Convolutional Network (TCN). By eliminating dependence on ER diagrams, domain-specific templates, and stable schemas, the approach enables generalizable and scalable reconstruction of execution traces in dynamic information systems. Experimental results on TPC-H/E, synthetic, and real-world industrial datasets demonstrate 85% accuracy in event prediction and recovery of approximately 82% of true predecessor relationships, yielding high-fidelity process traces.

execution behaviorOLTPprocess trace construction

Latest Papers

What's happening recently
View more

Existing database systems struggle to effectively model temporal dynamics, contextual dependencies, and causal relationships among attributes. To address this limitation, this work proposes Change Rules (CRs)—a novel rule-based paradigm that explicitly captures antecedent-consequent attribute changes within ordered tuple sequences, thereby overcoming the constraints of traditional data quality rules in modeling temporal and contextual patterns. The authors introduce CR-Miner, an efficient algorithm that employs a level-wise candidate generation strategy to identify change intervals, integrating declarative dependency specifications with sequence analysis techniques. Experimental results demonstrate that CR-Miner achieves a 40–50% average speedup over state-of-the-art methods while significantly enhancing the granularity and efficiency of trend analysis and causal inference.

Causal RelationshipsChange RulesData Profiling

This work addresses the challenge of silent updates to large language models (LLMs) by service providers, which often occur without version changes and can lead to behavioral drift and functional regressions, while existing mechanisms lack deployment-side control over compatibility governance. Framing LLM updates as a software supply chain governance problem, this study proposes a deployment-side control framework that defines rule-based production contracts, constructs risk-category-oriented test suites, and enforces compatibility gates to validate model safety and performance prior to updates. Experimental results demonstrate that the approach effectively uncovers fine-grained regressions missed by aggregate metrics, while also highlighting critical challenges in test design, threshold calibration, and drift attribution.

behavioral driftcompatibility governanceLarge Language Models

This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.

downstream tasksirrelevant messageslog analysis

Hot Scholars

RH

Robert Hirschfeld

Hasso Plattner Institute, University of Potsdam, Germany
Programming LanguagesSoftware ModularityExplorative ProgrammingComputational Reflection
CO

Chris Oehmen

Pacific Northwest National Laboratory
computational biologycybersecruity
AS

Arun Sharma

PhD Candidate Computer Science, University of Minnesota
Data MiningMachine LearningDatabase SystemsDistributed Systems
JM

Jason McDermott

Scientist Pacific Northwest National Laboratory
machine learninginfectious diseasestroke modelingbiological networks