Score
Designs and analyzes distributed synchronization protocols that implement gossip-based message exchanges and factor the synchronization process into multiple mixing phases, i.e., constructs algorithms that realize approximate or relaxed agreement via iterative mixing steps. Builds and evaluates protocol variants, convergence analyses, and parameter settings that trade agreement tightness for lower bandwidth and ensure graceful degradation under delays, message loss, and node failures.
Classical Lamport lower bounds for two-phase consensus in partially synchronous systems assume a minimum number of processes, yet practical protocols such as Egalitarian Paxos achieve two-phase decision-making with fewer processes—revealing an apparent contradiction rooted in the conflation of “consensus as an atomic object” and “consensus as a decision task.” Method: This paper formally distinguishes these two modeling paradigms and derives tight lower bounds for each under Byzantine (f) and crash (e) fault models. Contribution/Results: We establish that the minimal process requirement is max{2e + f − 1, 2f + 1} for consensus as an atomic object, and max{2e + f, 2f + 1} for consensus as a decision task. Both bounds are proven tight via constructive protocols that match them exactly. This resolves the theory–practice gap, provides the first precise and operationally meaningful characterization of the minimal process count for two-step consensus, and unifies foundational understanding across fault-tolerant distributed computing models.
This paper investigates the design of optimal consensus protocols for synchronous systems under crash-failure models. To address constrained information exchange, we propose a novel framework integrating failure counting with value association, achieving only a one-round increase in decision latency while significantly reducing computational overhead and storage requirements. Leveraging epistemic logic modeling and knowledge-based program implementation, we formally analyze and optimize FloodSet and its variants. Our work establishes, for the first time, strictly optimal consensus protocols for multiple classical information-exchange patterns—achieving performance asymptotically approaching the Dwork–Moses lower bound. Moreover, the protocols yield substantial improvements in both computational efficiency and space complexity over prior approaches.
研究在随机广播模型下,n个进程通过同步广播消息达成二元共识的问题,设计了在给定轮数内最小化错误分歧概率的算法。
Distributed computing theory lacks a unified pedagogical and research framework addressing fundamental challenges—including consensus impossibility (FLP), Byzantine fault tolerance, logical time, topological unsolvability, and self-stabilization—across synchronous and asynchronous models. Method: We systematically construct a formal, unified framework covering 36 core topics (e.g., consensus, fault tolerance, synchronization, shared memory, distributed graph algorithms), integrating advanced theoretical tools: BG-simulation, topological methods, obstruction-freedom, population protocols, formal object classification, failure detectors, quorum systems, randomized consensus, and fixed-point analysis. Contribution/Results: The framework establishes provability guarantees and bridges theory with practice via rigorously structured curricula. It has become the de facto standard for graduate-level distributed systems theory courses at numerous leading universities worldwide, significantly advancing the integration of mathematical rigor and engineering relevance in distributed systems education and research.
The FLP impossibility theorem asserts that deterministic binary consensus cannot simultaneously satisfy termination, agreement, and validity in a fully asynchronous system subject to crash failures. Method: This work challenges the universality of this conclusion by rigorously distinguishing “unattainable termination” from “undecidable protocol,” demonstrating that FLP only refutes strong agreement (i.e., all correct processes deciding the same value), not termination per se. We introduce a novel termination paradigm and design the first deterministic consensus algorithm that guarantees both termination and input-consistent decision values under the FLP model. Contribution/Results: Our algorithm converges—under any crash-failure pattern—to some initial input value, strictly satisfying termination, agreement, and validity. Theoretical correctness is established via distributed consensus theory, asynchronous state-machine replication, and the crash-stop failure model. This resolves a foundational limitation by showing deterministic consensus with guaranteed termination is achievable without relaxing asynchrony or fault assumptions.
Traditional asynchronous Byzantine Fault-Tolerant (BFT) atomic broadcast protocols incur excessive communication overhead—O(n²) per request—due to global broadcasting of all client requests, leading to redundant transmissions and resource waste. Method: This paper proposes Slim-HBBFT, a novel protocol centered on Priority-Provable Broadcast (P-PB), wherein only a designated subset of nodes generates lightweight broadcast proofs, reducing per-request communication complexity to O(n) while strictly satisfying the safety and liveness requirements of Asynchronous Common Subset (ACS). Slim-HBBFT integrates asynchronous BFT, provable broadcast, and priority-based scheduling, and its correctness is formally verified. Contribution/Results: Experimental evaluation demonstrates that Slim-HBBFT achieves significantly higher throughput and lower latency under high-concurrency workloads, establishing a new paradigm for efficient and secure distributed consensus in asynchronous settings.
This study addresses the realizability of global distributed protocols under asynchronous network architectures—specifically, whether local implementations can satisfy global specifications. To this end, the work introduces a network-parameterized coherence condition, combined with operational axioms that characterize message buffering behavior, enabling a unified formal model of five prominent asynchronous network paradigms. Leveraging symbolic algorithms and formal verification techniques, the paper establishes—for the first time—a systematic relationship between network architecture parameters and protocol realizability, and derives optimal complexity bounds. The accompanying tool, Sprout(A), is the first realizability verifier supporting multiple network architectures, achieving both high performance and modularity without sacrificing generality.
This work addresses the lack of efficient and practical vector consensus protocols in asynchronous networks, which has constrained the performance of Asynchronous Common Subset (ACS). The paper introduces an aggregated vector consensus primitive and, for the first time, applies it in a fully asynchronous setting to construct the JUNO protocol, achieving ACS with optimal $O(n^2)$ message complexity. Experimental evaluations demonstrate that JUNO improves average throughput by 93% over HoneyBadgerBFT and by 47% over Dory, thereby substantiating its high efficiency and practicality. This advancement effectively fills the gap in high-performance vector consensus protocols for such environments.
This work addresses the limitations imposed by the FLP impossibility result on deterministic consensus in asynchronous systems by proposing an event-synchronized vector consensus algorithm. It distinguishes between data-independent and data-dependent consensus, uncovering three implicit assumptions underlying the FLP theorem and demonstrating that a key assumption lacks empirical support. By integrating an event-driven synchronization mechanism with formal verification and experimental evaluation, the proposed protocol achieves both safety and liveness in a fault-tolerant manner. Experimental results show that the algorithm tolerates single-node crash failures while effectively transcending the practical applicability boundary of the FLP impossibility result.
This study addresses the inherent tension in consensus protocols, which are efficient under partial synchrony yet prone to stalling in asynchronous environments. To resolve this, we propose a hybrid consensus algorithm built upon a shared directed acyclic graph (DAG). The approach alternates between partially synchronous and asynchronous commit rules, dynamically adjusting execution periods to balance performance with liveness. It enables adaptive mode switching without additional communication overhead and incorporates a hidden leader mechanism to guarantee asynchronous liveness. The protocol is formally verified using Lean 4. Experimental results demonstrate that the proposed method achieves throughput comparable to synchronous protocols under favorable network conditions while preserving liveness approaching that of asynchronous protocols in adverse environments, thereby effectively reconciling efficiency with robustness.