Score
Designs and analyzes formal calibration guarantees and certification procedures for systems that use routing mechanisms (including hard/discrete routing and expert gating), proving relationships between expert-level and mixture-level calibration and establishing hard-routing calibration guarantees. Builds formal characterizations of how routing decisions interact with calibration—e.g., which distribution shifts preserve sufficiency—and develops methods to certify or test those properties.
Existing routing mechanisms struggle to reliably determine—prior to deployment—whether multi-agent large language model systems genuinely benefit from routing, as reliance solely on advisor complementarity or AUC often leads to erroneous conclusions. This work proposes RouteGuard, a novel framework that decouples routing gain from AUC for the first time by decomposing the gain into a policy improvement term π and a conditional error gap Δ_E, and introduces a conditional regret functional Φ to construct certifiable performance intervals under finite-sample settings. The method establishes sharp bounds matching the Le Cam lower bound, reveals a robustness phase transition phenomenon, and validates its efficacy through workload-aware clustering-based resampling and preregistered semi-synthetic experiments. Evaluated on RouterBench and OpenRCA benchmarks, RouteGuard accurately identifies effective routing scenarios, avoids spurious certification of redundant advisors, and demonstrates well-calibrated reliability.
Existing verification techniques lack end-to-end behavioral correctness guarantees for P4 data-plane programs deployed across programmable switches and smart NICs. Method: We propose the first formal verification framework comprehensively covering P4 control blocks, parsers/deparsers, and non-P4 hardware components—including multicast engines, resubmit paths, and packet generators. Our approach enables the first compositional correctness proof of P4 programs jointly with fixed or configurable hardware, transcending prior control-plane-only verification. It integrates formal semantics modeling, SMT-based reasoning, and modular specification to achieve fully automated, end-to-end data-plane verification. Results: We experimentally validate the complete packet-forwarding behavior of two canonical P4 applications, establishing system-level functional correctness. All verification results are mathematically provable and composable, ensuring rigorous, scalable assurance for heterogeneous P4-accelerated platforms.
This work addresses the critical challenge of safely replacing black-box models with low-cost, interpretable proxy models during deployment while ensuring controlled performance degradation. We propose an active routing mechanism that employs a lightweight gating model to decide at inference time whether to invoke the proxy, coupled with Clopper–Pearson conformal calibration to rigorously bound the rate of performance degradation violations without any distributional assumptions. To our knowledge, this is the first approach to provide distribution-free safety guarantees for model routing, establishing theoretical feasibility conditions and an AUC-based threshold criterion. We further show that probabilistic calibration primarily affects routing efficiency rather than validity. Experiments across 35 OpenML datasets demonstrate that our method substantially outperforms regression-based conformal and naive baselines, achieving significantly higher proxy usage coverage while strictly controlling violation rates.
This study addresses the evaluation bias and attribution challenges in multi-verifier routing caused by shifting contract conditions, formulating routing as a contract condition identification problem. Methodologically, it introduces core mechanisms including contract lattices, policy-agnostic response bands, and request-level bounds. By integrating causal inference, adaptive policy learning, and RLVR, the approach achieves precise attribution through matching path comparisons and generates verifiable certificates. Experimental results demonstrate that the proposed method effectively decouples policy gains from verifier set discrepancies, reducing attribution error to 0.0011 and significantly enhancing the reliability of system evaluation.
BGP configuration verification faces dual challenges of privacy leakage and scalability in large-scale networks. This paper introduces malicious-secure multi-party computation (MPC) to BGP verification for the first time, proposing an end-to-end privacy-preserving routing policy verification system. It combines formal BGP semantic modeling with a distributed convergence verification algorithm to enable zero-knowledge correctness proofs. The system achieves sub-3-second verification latency on thousand-node topologies and is formally proven to satisfy strong simulation-based privacy—guaranteeing that no participant learns any raw routing configurations or policies, even during cross-AS collaborative verification. Its core innovation lies in the tight integration of MPC protocols with BGP semantics, achieving both rigorous information-theoretic privacy and practical scalability for real-world deployment.
Traditional symbolic network verifiers rely on manual modeling, which struggles to accommodate vendor-specific implementation discrepancies and protocol evolution. This work proposes the first self-evolving verification framework that treats router software as a trusted oracle and integrates SMT solvers, counterexample-guided inductive synthesis (CEGIS), network emulation, and intelligent code agents within a closed-loop system to automatically extend and correct symbolic models. By shifting the verification focus from manual modeling to systematic testing, the approach autonomously uncovers vendor-specific behaviors. A prototype implementation successfully augmented a 3,000-line SMT-based verifier with support for OSPF areas, BGP route reflection, and L3VPN over EVPN, achieving complete behavioral consistency with the oracle.
This study addresses the lack of deductive reasoning foundations in discrete-event simulation, which hinders formal verification of model correctness and performance guarantees. To overcome this, we propose a core imperative calculus and proof system that extends reasoning over discrete-time probabilistic programs to performance models involving continuous time and distributions, rigorously establishing the soundness and completeness of the associated proof rules. Leveraging measure theory and the Lean theorem prover, we implement reasoning for almost-sure reachability and expected hitting times in continuous-time probabilistic programs. We successfully complete formal proofs on client-server architectures and network routing protocol case studies, thereby transcending the limitations of analytical solutions in classical queueing theory.
This study addresses the lack of distribution-free selective guarantees and low certificate reliability in chain-of-thought verifiers under small calibration budgets. To this end, it investigates abstention-based verifiers and proposes fixed-sequence certificates to achieve distribution-free guarantees in small-sample regimes. Methodologically, this work elucidates the principles underlying effective abstention and introduces certified lower bounds alongside Benjamini-Hochberg lattice conditions. By integrating conformal prediction, residual stream probing, and cross-fitting techniques, a novel certificate algorithm is developed that eliminates the need for monotonicity assumptions. Experimental results demonstrate that the proposed certificates consistently outperform Bonferroni methods in coverage across diverse model signals, significantly improving coverage levels for non-trivial targets.
This work addresses the longstanding divide between software and hardware verification, which has been hindered by the absence of a common intermediate representation that would enable direct application of efficient hardware model checking techniques to C programs. To bridge this gap, the paper introduces the Circuit-based Program Verification (CPV) framework, which systematically compiles C programs into sequential circuits, unifying control-flow and data-flow semantics within a single formal model. CPV integrates established hardware model checking algorithms—including Bounded Model Checking (BMC), k-induction, and IC3/PDR—to support both reachability safety and termination verification. Moreover, it automatically translates counterexamples back into human-readable software evidence. Evaluated on a benchmark suite of over 16,000 verification tasks, CPV matches the performance of leading software verifiers and successfully solves instances beyond the reach of existing tools, demonstrating significant complementary strengths.
本文解决了电路最小化中无法验证最优性声明的问题,通过合成GF(2)上的最小线性程序并生成可验证的DRAT证明来确保声明的两部分都得到认证。