implement append-only logging

Designs and builds storage and software components that record immutable, sequential records of events or state changes as append-only writes, including schemas, write paths, and data encodings that prevent in-place mutation. Implements supporting features such as durable write and recovery guarantees, indexing/compaction and retention strategies, access controls and APIs for audit, replay, and verification of the event stream.

implementappend-onlylogging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.86
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the problem of providing provably correct snapshot-equivalent change capture and replay for continuously written databases without relying on global snapshots or locking mechanisms. It introduces the “authenticated virtual slice” model, which combines interleaved log scanning with watermark-based positioning to enable incremental authentication over primary key ranges and advancing frontiers while preserving source continuity. For the first time, the paper formally defines and machine-verifies snapshot equivalence for DBLog, presenting a rigorous Isabelle/HOL proof that all well-formed executions satisfy per-key replay equivalence within their designated frontiers and key ranges, and that the authentication process yields valid virtual slices.

change-data-capturecorrectness formalizationdatabase replication

This work addresses reference instability in tiered append-only sequences—manifesting as dangling, stale, corrupted references, and snapshot skew—caused by migrating hot data to cold tiers. To resolve this, the authors propose the Cascade Log structure, which formally defines cross-tier reference anomalies for the first time and introduces a persistent merge-interval mapping as the authoritative version source. By integrating immutable root snapshots, block-level folding, and B-tree-like update strategies, Cascade Log achieves both reference stability and snapshot consistency under append-heavy workloads. The resulting index structure exhibits Θ(A) space complexity, supports point and range queries in O(log A) and O(log A + k) time respectively, and attains theoretically optimal update overhead. Experimental results demonstrate that the approach effectively eliminates reference anomalies at scale, even with millions of records.

append-only logscross-tier anomaliesreference stability

This work addresses the persistent degradation of agent reasoning and tool use caused by erroneous memories—such as contamination, staleness, or misattribution—which existing approaches struggle to correct without discarding valid knowledge. The paper formalizes, for the first time, the post-failure memory recovery problem and introduces a dependency-guided rollback repair mechanism. By constructing a typed memory-action dependency graph, the method tracks downstream effects at runtime, selectively deactivates unreliable memories, and replays only those computations relevant to the final answer. Evaluated on a controlled benchmark of 150 cases, the approach achieves an 85.3% recovery rate—surpassing the best baseline (77.3%)—while fully eliminating error sources and preserving all benign memories. In 50 stress-test scenarios, it attains a 68.0% recovery rate and significantly outperforms baselines, achieving the highest statement invalidation F1 score of 0.669.

faulty memoriesmemory errorsmemory-augmented agents

Latest Papers

What's happening recently
View more

Existing non-volatile memory (NVM)-based key-value stores commonly lack support for modern database interfaces such as snapshots, consistent iterators, and atomic batch operations. To address this limitation, this work proposes and implements FlintKV, a skip-list-based NVM-optimized storage engine that natively supports a full production-grade transactional API while guaranteeing persistent linearizability. FlintKV achieves this through a novel concurrency protocol that integrates multi-version concurrency control with flat combining, leveraging byte-addressable NVM, custom persistence mechanisms, and highly efficient data structures. Experimental results demonstrate that FlintKV significantly enhances performance, achieving up to 75% higher end-to-end throughput compared to state-of-the-art alternatives. The system can be deployed standalone or integrated as a functional enhancement module into existing NVM-based storage systems.

atomic batchdurable linearizabilitykey-value store

This work addresses the challenge of knowledge updating in large language models, which typically necessitates costly retraining. Building upon the Compositional Multi-layer Memory (CMM) architecture, the authors propose a version-aware operational layer that compiles high-level semantic edits into ordered, composable memory primitive transactions. By introducing versioned CMM and transactional CMM, knowledge modifications are modeled as reversible and reusable structured transactions, enabling fine-grained replacement, rollback, historical tracing, and localized updates. This approach substantially reduces reliance on full model retraining while ensuring editing efficiency and traceability.

knowledge updatingmemory editingmulti-layer MeMo

This study addresses the vulnerability of LLM agents to task failures and unsafe operations caused by the reuse of invalid persistent states. To mitigate this, we propose a pre-action state diagnosis and repair framework that employs record-level counterfactual replanning to localize critical corrupted records, combined with typed evidence binding and reliability checks to repair states while maintaining lineage. Furthermore, independent verification is enforced prior to execution through read-only fact validation and intent clarification mechanisms. This approach enables continuous correction and decoupled state-action verification. Experimental results demonstrate that our method achieves 93.3% accuracy under corrupted states—substantially outperforming the 38.7% baseline—while ensuring zero unsafe actions and exhibiting strong cross-environment transferability.

coding agentsinvalid recordsLLM agents

This work addresses the challenge that flat patches generated by code agents lack structured commit histories, hindering code review, rollback, and maintenance. It formalizes retrospective commit history reconstruction as a code block segmentation task under replay constraints, introduces a benchmark dataset comprising 800 real sequential commits, and proposes multidimensional evaluation metrics—PPAR, ARI, and TCR. The method leverages large language models (e.g., GPT-5.4, GLM-5) augmented with code roles and dependency information for clustering, incorporating Dependency-Aware Clustering Evidence (DACE) to improve grouping accuracy and validating executability through replay. The best-performing model achieves an ARI of 0.46, significantly outperforming baselines; DACE further boosts low-scoring systems by 0.05–0.08 ARI, revealing that reconstructing real-world commit histories is substantially more difficult than synthetic scenarios.

atomic commitscode evolutioncommit history reconstruction

Hot Scholars

YW

Yu Wang

University of Science and Technology of China
LLM ReasoningLLM AgentReinforcement Learning
YZ

Yan Zhou

Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)
Machine TranslationSpeech Translation
YW

Yusong Wang

Tokyo Institute of Technology
Representation LearningAffective Computing
LL

Liuzhenghao Lv

Phd student Computer Science, Peking University
Large Language ModelsAI for ScienceSpiking Neural Networks
JS

Jingwei Sun

University of Science and Technology of China
High Performance ComputingPerformance ModelingArchitecture SimulationEfficient Deep Learning