Score
Design view-change protocols: design, build, and analyze the procedures that transition a replicated distributed consensus system between leaders or views, specifying pacemaker advancement rules, calibrating non-critical timeouts, defining quorum and safety conditions (including operation without strong quorums), and handling view-change failure modes so that safety and liveness are preserved.
This work addresses the performance degradation in parallel Byzantine Fault Tolerance (BFT) protocols caused by inefficient or unavailable node selection during view changes triggered by primary node failures. To tackle this challenge, the paper introduces, for the first time, a mixed-integer programming approach to optimize view change operations in parallel BFT, proposing the View Change Optimization (VCO) model. VCO jointly optimizes primary replica selection and follower reassignment by incorporating communication latency and fault scenarios. An efficient iterative algorithm for backup primary selection is developed using an enhanced decomposition method combined with Bendersโ cutting-plane technique. Experimental evaluation on Microsoft Azure demonstrates that VCO-driven parallel BFT significantly outperforms existing approaches under both normal and faulty conditions, with performance gains amplifying as network scale increases.
This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.
This study addresses the lack of a unified evaluation framework for the diverse array of blockchain consensus protocols, which hinders their effective deployment across applications due to inconsistent assessments of security, scalability, and energy efficiency. To bridge this gap, the paper proposes a systematic analytical framework that integrates literature review, protocol workflow modeling, and qualitative comparative analysis to holistically evaluate the operational mechanisms, applicability boundaries, and real-world performance of mainstream consensus algorithms in domains such as finance, supply chain, and healthcare. The work not only establishes a multidimensional evaluation methodology to guide protocol selection for both academia and industry but also identifies critical research gaps, thereby outlining key directions for future advancements in enhancing security, scalability, and energy efficiency.
Traditional consensus protocols rely on deterministic f-threshold fault models that struggle to capture the complex failure behaviors observed in real-world systems, thereby limiting optimization of performance and cost. This work proposes a novel consensus mechanism grounded in a probabilistic fault model, which incorporates machine-level failure curves and abandons the rigid majority quorum constraint in favor of dynamic, non-traditional quorum strategies. By more accurately reflecting actual operating conditions, the proposed approach substantially enhances system reliability, efficiency, cost-effectiveness, and sustainability.
This work addresses the challenge of achieving both liveness and strong safety in consensus protocols under conditions of dynamic participant availability and network asynchrony. The authors propose Majorum, a novel consensus architecture that integrates a quorum-based TOB-SVD protocol with a partially synchronous finality mechanism to realize an ebb-and-flow consensus model. This design enables participants to freely go offline and reconnect while ensuring strong safety through prefix immutability. Under optimistic network conditions, Majorum achieves efficient finality with single-round voting and confirms one block every three slots, substantially improving system throughput and responsiveness without compromising security guarantees.
This work addresses the loss of liveness in traditional Byzantine fault-tolerant protocols during network partitions, which arises from their reliance on strong quorums. The authors propose a two-layer authentication framework that decouples consensus availability from commitment, enabling each partition to independently extend its local chain while ensuring safe reconciliation upon network recovery. To support progress during partitions without requiring strong quorums, they introduce a view synchronization pacemaker that integrates proof-of-authority (PoA)-based speculative execution with timeout calibration along non-critical paths. Experimental results demonstrate that the system achieves up to 900,000 TPS with 16 nodes and 480,000 TPS with 104 nodes under stable network conditions (latency: 0.31โ0.75 seconds), while sustaining non-zero speculative throughput during severe partitionsโall speculative work remaining safely reconcilable afterward.
This work addresses the tight coupling between reconfiguration mechanisms and consensus protocols in existing state machine replication (SMR) systems, which hinders independent component upgrades and incurs significant downtime. To overcome this limitation, the authors propose Gauss, a reconfiguration engine that decouples the internal consensus log from an externally exposed clean log through a dual-log architecture. This separation enables modular evolution of membership changes, fault-tolerance thresholds, and even the underlying consensus protocol itself. By isolating reconfiguration logic from consensus execution, Gauss substantially reduces system maintenance complexity and supports seamless, independent upgrades of individual components. Evaluation on the Rialto blockchain platform demonstrates that Gauss facilitates near-zero-downtime transitions between different consensus protocols, achieving highly available and adaptable SMR deployments.
Current approaches to automated program synthesis lack effective governance mechanisms to ensure the compliance of generated code. This work proposes Protocol-Driven Development (PDD), a model that treats machine-executable protocols as primary artifacts and delineates the space of valid implementations through structural, behavioral, and operational invariants. PDD mandates that every implementation be accompanied by a verifiable chain of compliance evidence. By integrating formal methods, property-based testing, policy-as-code, and software provenance techniques, PDD establishes a unified framework for protocol specification and verification. This framework enables trustworthy admission control over automatically synthesized code, guaranteeing that all adopted implementations strictly adhere to protocol constraints and are backed by complete, auditable proofs of compliance.
This work addresses the challenge in model-based engineering where interdisciplinary teams struggle to comprehend each otherโs modifications due to change descriptions being tightly coupled with domain-specific modeling knowledge. To overcome this, we propose a method that automatically maps formal Delta Modeling changes into model-agnostic natural language descriptions, preserving semantic fidelity while enhancing human readability and cross-domain applicability. We introduce, for the first time, a systematic mapping mechanism between formal change representations and generic natural language. The approach is rigorously evaluated through an integrated framework comprising a prototype implementation, case studies, and user studies, demonstrating its technical feasibility, practical utility, and scalability. Our results establish a viable pathway toward automated generation of interpretable, human-readable descriptions of model changes.
This work addresses the challenge of ensuring safety, reliability, and trustworthiness in collective adaptive systems operating in dynamic environments by proposing a modular design paradigm centered on intrinsic trustworthiness. The approach integrates a runtime model based on local causal event sequences, a temporal logic verification technique supporting modular architectures, and a compositional reasoning mechanism for global system properties grounded in component attributes. Through this tripartite framework, the study overcomes key limitations of conventional formal methods and demonstrates substantial improvements in verifiability and scalability in case studies, thereby establishing both a theoretical foundation and a practical pathway for engineering highly trustworthy collective adaptive systems.