Score
Designs and builds storage and software components that record immutable, sequential records of events or state changes as append-only writes, including schemas, write paths, and data encodings that prevent in-place mutation. Implements supporting features such as durable write and recovery guarantees, indexing/compaction and retention strategies, access controls and APIs for audit, replay, and verification of the event stream.
研究解决了数据库行与变更日志交织时的复制到日志交接问题,通过提出并验证Generalized DBLog协议来确保数据一致性。
This work addresses the problem of providing provably correct snapshot-equivalent change capture and replay for continuously written databases without relying on global snapshots or locking mechanisms. It introduces the “authenticated virtual slice” model, which combines interleaved log scanning with watermark-based positioning to enable incremental authentication over primary key ranges and advancing frontiers while preserving source continuity. For the first time, the paper formally defines and machine-verifies snapshot equivalence for DBLog, presenting a rigorous Isabelle/HOL proof that all well-formed executions satisfy per-key replay equivalence within their designated frontiers and key ranges, and that the authentication process yields valid virtual slices.
This work addresses reference instability in tiered append-only sequences—manifesting as dangling, stale, corrupted references, and snapshot skew—caused by migrating hot data to cold tiers. To resolve this, the authors propose the Cascade Log structure, which formally defines cross-tier reference anomalies for the first time and introduces a persistent merge-interval mapping as the authoritative version source. By integrating immutable root snapshots, block-level folding, and B-tree-like update strategies, Cascade Log achieves both reference stability and snapshot consistency under append-heavy workloads. The resulting index structure exhibits Θ(A) space complexity, supports point and range queries in O(log A) and O(log A + k) time respectively, and attains theoretically optimal update overhead. Experimental results demonstrate that the approach effectively eliminates reference anomalies at scale, even with millions of records.
This work addresses the persistent degradation of agent reasoning and tool use caused by erroneous memories—such as contamination, staleness, or misattribution—which existing approaches struggle to correct without discarding valid knowledge. The paper formalizes, for the first time, the post-failure memory recovery problem and introduces a dependency-guided rollback repair mechanism. By constructing a typed memory-action dependency graph, the method tracks downstream effects at runtime, selectively deactivates unreliable memories, and replays only those computations relevant to the final answer. Evaluated on a controlled benchmark of 150 cases, the approach achieves an 85.3% recovery rate—surpassing the best baseline (77.3%)—while fully eliminating error sources and preserving all benign memories. In 50 stress-test scenarios, it attains a 68.0% recovery rate and significantly outperforms baselines, achieving the highest statement invalidation F1 score of 0.669.
论文探讨了如何从记忆中精确删除特定记录的问题,通过比较不同方法的效果,发现回放检查点是唯一实现精确删除的方法。
Existing non-volatile memory (NVM)-based key-value stores commonly lack support for modern database interfaces such as snapshots, consistent iterators, and atomic batch operations. To address this limitation, this work proposes and implements FlintKV, a skip-list-based NVM-optimized storage engine that natively supports a full production-grade transactional API while guaranteeing persistent linearizability. FlintKV achieves this through a novel concurrency protocol that integrates multi-version concurrency control with flat combining, leveraging byte-addressable NVM, custom persistence mechanisms, and highly efficient data structures. Experimental results demonstrate that FlintKV significantly enhances performance, achieving up to 75% higher end-to-end throughput compared to state-of-the-art alternatives. The system can be deployed standalone or integrated as a functional enhancement module into existing NVM-based storage systems.
This work addresses the challenge of knowledge updating in large language models, which typically necessitates costly retraining. Building upon the Compositional Multi-layer Memory (CMM) architecture, the authors propose a version-aware operational layer that compiles high-level semantic edits into ordered, composable memory primitive transactions. By introducing versioned CMM and transactional CMM, knowledge modifications are modeled as reversible and reusable structured transactions, enabling fine-grained replacement, rollback, historical tracing, and localized updates. This approach substantially reduces reliance on full model retraining while ensuring editing efficiency and traceability.
This study addresses the vulnerability of LLM agents to task failures and unsafe operations caused by the reuse of invalid persistent states. To mitigate this, we propose a pre-action state diagnosis and repair framework that employs record-level counterfactual replanning to localize critical corrupted records, combined with typed evidence binding and reliability checks to repair states while maintaining lineage. Furthermore, independent verification is enforced prior to execution through read-only fact validation and intent clarification mechanisms. This approach enables continuous correction and decoupled state-action verification. Experimental results demonstrate that our method achieves 93.3% accuracy under corrupted states—substantially outperforming the 38.7% baseline—while ensuring zero unsafe actions and exhibiting strong cross-environment transferability.
This work addresses the challenge that flat patches generated by code agents lack structured commit histories, hindering code review, rollback, and maintenance. It formalizes retrospective commit history reconstruction as a code block segmentation task under replay constraints, introduces a benchmark dataset comprising 800 real sequential commits, and proposes multidimensional evaluation metrics—PPAR, ARI, and TCR. The method leverages large language models (e.g., GPT-5.4, GLM-5) augmented with code roles and dependency information for clustering, incorporating Dependency-Aware Clustering Evidence (DACE) to improve grouping accuracy and validating executability through replay. The best-performing model achieves an ARI of 0.46, significantly outperforming baselines; DACE further boosts low-scoring systems by 0.05–0.08 ARI, revealing that reconstructing real-world commit histories is substantially more difficult than synthetic scenarios.
本文提出了一种基于对等复制的持久工作流状态架构,通过认证的追加日志和显式证书来解决现有框架中状态权威性和信任问题。