design fault-tolerant protocols

Designs, builds, and analyzes distributed protocols that maintain correctness (safety) and progress (liveness) despite faults and adversarial (Byzantine) behavior, including modeling adversaries and failure thresholds and proving safety and liveness bounds. Work in this skill also includes implementing Byzantine-resilient protocol steps and constructing and analyzing fault-tolerant routing mechanisms and other fault-tolerance analyses so systems continue to operate under crashes, message loss, or malicious nodes.

designfault-tolerantprotocols

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Byzantine Fault Tolerance Approach towards AI Safety

Apr 20, 2025
JD
John deVadoss
🏛️ Global Blockchain Business Council | Deutsche Bank

This work addresses the reliability of AI systems under unexpected failures and adversarial attacks. We systematically adapt Byzantine Fault Tolerance (BFT)—a foundational paradigm from distributed systems—to AI safety, introducing the novel conceptual analogy that “malicious AI modules correspond to Byzantine nodes.” Based on this, we formalize a component-level AI failure model and design a multi-agent consensus verification framework integrating distributed consensus protocols, behavioral consistency checking, redundant heterogeneous model arbitration, and controlled fault-injection testing. Evaluated across multiple high-stakes AI decision-making tasks, our architecture achieves a 99.2% anomaly detection rate, substantially enhancing robustness and trustworthiness against both adversarial perturbations and internal component failures. The proposed approach establishes a verifiable, scalable, and principled new paradigm for AI safety.

Applying BFT to AI safetyEnsuring AI reliability under faultsUsing consensus for safe AI

Formal verification of liveness in Byzantine Fault-Tolerant (BFT) distributed systems remains challenging due to the difficulty of rigorously modeling adversarial behavior, cryptographic primitives, and progress guarantees. Method: This paper introduces three novel, modular techniques grounded in formal verification: (i) compositional proof support for sub-protocol separation; (ii) precise semantic modeling of cryptographic signatures; and (iii) joint verification of liveness and safety properties. The framework is deeply integrated into an executable Go implementation of Practical Byzantine Fault Tolerance (PBFT). Contribution/Results: It achieves the first end-to-end formal liveness verification of a single-log PBFT consensus protocol. Experimental evaluation confirms that the prototype strictly guarantees both progress and correctness under normal operation and diverse malicious fault scenarios. This work establishes the first practical, formally verified path to liveness assurance in BFT systems.

Ensuring PBFT correctness and liveness in executable implementationsModular verification for distributed systems with malicious participantsProving liveness in Byzantine decentralized systems

Accountable Liveness

Apr 16, 2025
AL
Andrew Lewis-Pye
🏛️ London School of Economics | a16z Crypto Research | Columbia University | Ethereum Foundation

Consensus protocols often lack accountability guarantees for liveness—i.e., the ability to uniquely identify and provably attribute liveness violations to a majority of malicious nodes. Method: We formally define *liveness accountability*: when liveness fails, at least a strict majority of Byzantine nodes must be uniquely identifiable and provably culpable. To capture realistic network uncertainty, we introduce the *x-partial synchrony* model, which unifies asynchronous and synchronous behaviors via a tunable parameter *x*. Within this model, we rigorously characterize the necessary and sufficient conditions for accountable liveness: *x < 1/2* and *f < n/2*, where *f* is the number of Byzantine nodes and *n* the total number of nodes—thereby establishing its fundamental feasibility boundary. Contribution/Results: We design a near-optimal protocol achieving asymptotically optimal culpable-node identification. Our work provides the first formal foundation and optimality proof for mechanisms such as Ethereum’s “inactivity leak,” bridging theory and practice in accountable consensus.

Achieving accountable liveness in consensus protocolsCharacterizing conditions for x-partially-synchronous networksOptimizing violation identification in liveness failures

A New Probabilistic Mobile Byzantine Failure Model for Self-Protecting Systems

Nov 06, 2025
SB
Silvia Bonomi
🏛️ Sapienza University of Rome | Niccolò Cusano University | Technion | Sorbonne University

To address the challenges posed by dynamically evolving, cross-layer attacks in modern distributed systems, this paper proposes a Probabilistic Mobile Byzantine Fault (P-MBF) model that transcends traditional static or deterministic Byzantine assumptions. The model explicitly captures the stochastic propagation of attacks across nodes and the system’s autonomous recovery as a coupled process. Integrated into the analysis component of the MAPE-K autonomic architecture, it enables real-time security situational assessment and policy-driven dynamic reconfiguration. We theoretically derive expected time thresholds for system transitions to safe or hazardous states; stochastic modeling and simulation validate the model’s sensitivity and effectiveness under varying infection and recovery rates. Our key contribution is the first integration of probabilistic mobile fault modeling with closed-loop self-protection control, establishing a quantifiable and verifiable paradigm for resilient system design in dynamic adversarial environments.

Analyzes Byzantine node threshold crossing and system self-recovery timingModels attack spread versus recovery rates in MAPE-K distributed systemsProposes probabilistic Mobile Byzantine Failure model for evolving attacks

Verification and Attack Synthesis for Network Protocols

Nov 02, 2025
MV
Max von Hippel
🏛️ Northeastern University

Ensuring functional correctness and performance resilience of network protocols under component failures and adversarial attacks remains a significant challenge. Method: This paper proposes a synergistic analysis framework integrating formal verification with attack synthesis. It models protocol behavior using a formal specification language and employs logical predicates, trace analysis, and model checking to achieve closed-loop verification—simultaneously establishing correctness guarantees and automatically generating realistic attack scenarios. Contribution/Results: Diverging from conventional unidirectional verification, our approach innovatively embeds attack-path generation directly into the verification workflow, enabling reproducible and interpretable failure attribution. Experimental evaluation across multiple mainstream network protocols demonstrates substantial improvements in vulnerability detection rates and attack-surface characterization accuracy. The results validate the feasibility and practicality of formal methods for deep, security-critical analysis of complex network protocols.

Applying formal methods to analyze protocols under normal and attack conditionsSynthesizing attacks that prevent protocol requirement achievementVerifying network protocol functionality and performance requirements

Latest Papers

What's happening recently
View more

This work investigates the fault-tolerance limits and protocol design for low-latency consensus under a hybrid failure model combining Byzantine faults (f) and crash faults (c). It establishes, for the first time, a tight lower bound of n ≥ 5f + 3c + 1 for two-message-latency commit protocols. The paper proposes a hybrid fault-tolerant consensus protocol featuring both a fast-commit path and a resilient fallback mechanism, enabling clients to select their desired finality latency. Built upon the partial synchrony model and integrating multi-round safety paths with synchronous recovery, the protocol achieves high performance: under a configuration of n=99, f=16, and c=6, it tolerates up to 22% failed replicas (liveness), 16% malicious nodes with 1-RTT safety, and as many as 54% malicious nodes with 2-RTT safety.

Byzantine faultscrash faultsfault tolerance

Existing large language model (LLM) approaches struggle to detect deep logical flaws in consensus protocols arising from multi-stage, complex state dependencies, often leading to violations of safety properties. This work proposes Agora, a domain-aware multi-agent verification framework that introduces, for the first time, a role-specialized multi-agent collaboration mechanism. By integrating hypothesis-driven testing, domain-constrained state-space exploration, and iterative validation, Agora enables systematic reasoning about global protocol invariants, overcoming the limitations of conventional single-function code analysis. Evaluated on four prominent consensus protocols—Raft, EPaxos, HotStuff, and BullShark—Agora successfully uncovers 15 previously unknown safety violations, all of which eluded detection by current LLM-based agents.

consensus protocolsdistributed systemsprotocol-level logic bugs

This work addresses the state explosion problem inherent in asynchronous, parameterized distributed protocols, which arises from communication asynchrony and unbounded participant counts. The authors propose an automated safety verification method based on backward unreachableness analysis. Their key innovation lies in distinguishing parameterized unboundedness into affine and non-affine categories, focusing specifically on affine protocols. By integrating goal-directed instantiation, causal reasoning, and state summarization, the approach efficiently prunes the state space. The prototype tool DissProve successfully verifies multiple affine protocols featuring infinitely many participants and unbounded execution lengths, achieving—for the first time—scalable, fully automatic safety verification for such asynchronous parameterized systems.

asynchronous systemsautomated verificationdistributed protocols

This work addresses a critical limitation in traditional Byzantine fault tolerance (BFT), which assumes honest nodes correctly enforce protocol semantics—an assumption that fails in agent-based systems where compliant nodes may erroneously endorse semantically invalid state transitions due to reasoning errors, thereby compromising execution safety. To resolve this, the paper introduces Epistemic Byzantine Fault Tolerance (EBFT), formally defining "epistemic faults" and the "honest majority problem," thereby decoupling semantic correctness from mere protocol compliance. The authors develop a two-dimensional fault-tolerance framework using confidence-weighted parameters: \(e_\delta\) for semantic safety risk and \(u_\varepsilon\) for liveness degradation. By integrating probabilistic belief modeling with tail concentration analysis, they derive novel quorum conditions that jointly guarantee semantic validity, consensus consistency, and system liveness. The analysis shows that fault tolerance improves only when newly added agents substantially reduce the tail risks of invalid endorsements or unavailable support.

Agentic InfrastructureByzantine Fault ToleranceEpistemic Fault

Hot Scholars

JD

Jérémie Decouchant

Assistant Professor (Universitair Docent, tenured), Delft University of Technology
ResilienceDistributed Systems & AlgorithmsDistributed Ledger TechnologyDistributed Learning
JC

Jinyuan Chen

Ocior Inc. & Louisiana Tech University
Information TheoryDistributed Consensus