Score
Designs, builds, and evaluates systems that detect and extract modifications made to persistent data stores and emit those modifications as ordered change events or logs for downstream consumers. This includes engineering capture mechanisms, schema-evolution handling, delivery semantics (ordering, latency, and at-least/at-most/exactly-once guarantees), and integration with replication, streaming pipelines, or event-sourced materializations.
研究解决了数据库行与变更日志交织时的复制到日志交接问题,通过提出并验证Generalized DBLog协议来确保数据一致性。
Databases continuously evolve through operations such as schema changes, version updates, and data transformations; however, existing approaches typically address these functionalities in isolation, lacking a unified abstraction. This work proposes the first integrated model that unifies continuous schema evolution, version management, and data transformation within a single framework. Built upon general-purpose computational primitives, the model supports operation provenance, conditional update propagation, and change alerts, while employing a declarative mechanism to manage the co-evolution of dependent artifacts—including views and machine learning models. A prototype system implements this framework using an enhanced, parameterized Prolly Tree—a Merkle tree–inspired data structure—to construct a relational-like engine. Experimental evaluation demonstrates that the proposed approach is both feasible and offers tunable performance across diverse evolution scenarios.
This paper addresses version control challenges for multidimensional structured data—namely, temporal evolution, spatial collaboration, and design iteration—by introducing Operational Differencing, a novel paradigm. Methodologically, it incorporates high-level semantic operations (e.g., schema changes, refactorings) into the version model; adopts an append-only branch history with a repository-free, lightweight “copy-as-branch” architecture; and enables operational query representation and future-tense execution. Contributions include: (1) the first systematic support for automatic schema-adaptive query rewriting under schema evolution; (2) precise, fine-grained diff/merge across structural transformations; (3) resolution of four out of eight canonical schema evolution challenges; and (4) a simplified versioning experience with no explicit repository and asymptotically zero branching overhead.
This work addresses the problem of providing provably correct snapshot-equivalent change capture and replay for continuously written databases without relying on global snapshots or locking mechanisms. It introduces the “authenticated virtual slice” model, which combines interleaved log scanning with watermark-based positioning to enable incremental authentication over primary key ranges and advancing frontiers while preserving source continuity. For the first time, the paper formally defines and machine-verifies snapshot equivalence for DBLog, presenting a rigorous Isabelle/HOL proof that all well-formed executions satisfy per-key replay equivalence within their designated frontiers and key ranges, and that the authentication process yields valid virtual slices.
Modern OLTP systems often suffer from frequent schema changes, missing primary/foreign keys, and fragmented execution traces, rendering traditional approaches—reliant on fixed schemas and manual modeling—costly and error-prone. This work proposes a fully automated pipeline that operates without predefined schemas by identifying quasi-key and timestamp columns, discovering inter-table relationships through statistical signals, and assembling and ordering events accordingly. To capture long-range dependencies across system events, the method incorporates a Temporal Convolutional Network (TCN). By eliminating dependence on ER diagrams, domain-specific templates, and stable schemas, the approach enables generalizable and scalable reconstruction of execution traces in dynamic information systems. Experimental results on TPC-H/E, synthetic, and real-world industrial datasets demonstrate 85% accuracy in event prediction and recovery of approximately 82% of true predecessor relationships, yielding high-fidelity process traces.
Existing database systems struggle to effectively model temporal dynamics, contextual dependencies, and causal relationships among attributes. To address this limitation, this work proposes Change Rules (CRs)—a novel rule-based paradigm that explicitly captures antecedent-consequent attribute changes within ordered tuple sequences, thereby overcoming the constraints of traditional data quality rules in modeling temporal and contextual patterns. The authors introduce CR-Miner, an efficient algorithm that employs a level-wise candidate generation strategy to identify change intervals, integrating declarative dependency specifications with sequence analysis techniques. Experimental results demonstrate that CR-Miner achieves a 40–50% average speedup over state-of-the-art methods while significantly enhancing the granularity and efficiency of trend analysis and causal inference.
本文提出了一种基于来源感知的确定性优先方法,从Oracle DDL重构逻辑和概念数据规范,以解决遗留数据库迁移中因文档不全导致的问题。
为了解决程序修复过程中需求到修复过程的可审查性问题,提出THEMIS方法,通过语义解析、需求-代码图等手段实现过程外部化。
This work addresses the challenge of silent updates to large language models (LLMs) by service providers, which often occur without version changes and can lead to behavioral drift and functional regressions, while existing mechanisms lack deployment-side control over compatibility governance. Framing LLM updates as a software supply chain governance problem, this study proposes a deployment-side control framework that defines rule-based production contracts, constructs risk-category-oriented test suites, and enforces compatibility gates to validate model safety and performance prior to updates. Experimental results demonstrate that the approach effectively uncovers fine-grained regressions missed by aggregate metrics, while also highlighting critical challenges in test design, threshold calibration, and drift attribution.
This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.