Machine-Checked Dual-Write Recovery from a Committed Log

📅 2026-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of achieving exactly-once semantics in distributed systems recovering from crashes under a shared-nothing architecture. By formally modeling the information-theoretic boundaries inherent in crash recovery, it demonstrates that relying solely on locally persisted state inevitably leads to either duplicate or missed deliveries. The paper presents the first rigorous, machine-checkable definition of exactly-once semantics together with precise conditions for its realizability. Leveraging a formal framework built in Isabelle/HOL that integrates information-theoretic limits, crash timing, receiver-side fencing, and evidence validity analysis, the study proves that conventional dual-write protocols suffer from fundamental uncertainty. It then introduces a provably safe recovery mechanism and quantifies how memory constraints for deduplication and source log truncation impact the guaranteed validity window.
📝 Abstract
After a crash, a delivery process faces a question its own database cannot answer: did the other side already receive the effect? Transactional outboxes and change data capture remove the application's dual write, but the relay they introduce delivers and records its own progress as two separate durable acts, so the same decision reappears one stage later. Practitioners have handled this boundary for a decade with retries, checkpoints, idempotency keys, and fencing, and the operational advice is largely sound. What has been missing is a precise account of when it works: the exact event the guarantees refer to, the evidence they require, and how long that evidence lasts. This paper supplies the missing account as a machine-checked theory, developed in Isabelle/HOL. At its heart is an information bound. Two reachable post-crash states can agree on everything the crashed side durably knows and still differ in what the sink accepted, so any recovery decision computed from that side must duplicate a delivery or leave one owed. This is not a story about sloppy bookkeeping: a single deterministic deliver-then-checkpoint protocol, its own durable cursor included in what recovery reads, is defeated by crash timing alone. Reading the sink's accepted record escapes the bound exactly, under stated premises. The answer can then go stale -- an old request still in flight, a second recoverer racing the first -- and each hazard has its own proved fence at the sink's acceptance boundary, the cost of fencing itself a theorem. Finally, the guarantee has a lifetime: bounded deduplication memory and truncated source history each void it in a proved way. The resulting test for any exactly-once recovery claim is short: what did the sink accept, what can still change that answer, and how long will the evidence survive?
Problem

Research questions and friction points this paper is trying to address.

exactly-once delivery
crash recovery
dual-write problem
durability
idempotency
Innovation

Methods, ideas, or system contributions that make the work stand out.

dual-write recovery
machine-checked proof
exactly-once delivery
information bound
fencing
🔎 Similar Papers
No similar papers found.