Score
Designs, implements, and analyzes test suites and automated testing infrastructure to verify the numerical correctness of algorithms, simulations, libraries, or software components that perform arithmetic, floating‑point, or other numeric computation. This includes building unit and regression tests, property- and example-based tests, reference-implementation comparisons and diagnostics to detect and quantify rounding error, convergence, stability, sensitivity, reproducibility, and violations of specified numerical tolerances.
FloatLib使用Lean 4开发了一个验证过的任意精度浮点数算术库,统一了多种格式和舍入规则,通过验证的软件后端保证正确性和效率。
This work addresses the problem of deriving provably tight floating-point rounding error bounds for numerical programs featuring conditional branches, no loops, and mixed-precision arithmetic. Methodologically, it unifies the modeling of conditional control flow and precision heterogeneity via two novel quantitative metrics—“instability jumps” and “window width”—and integrates interval arithmetic, abstract interpretation, and precision-aware semantic modeling, augmented with abstraction-guided global optimization. Its key contribution is the first formal framework enabling joint, compositional analysis of conditional branching and mixed precision, achieving both high bound tightness and practical analysis efficiency. Experimental evaluation on standard benchmarks demonstrates significantly tighter error bounds compared to prior approaches. Furthermore, the framework successfully guides precision configuration—e.g., step size and search direction—in the conjugate gradient method, empirically validating its utility in supporting design-time trade-offs among accuracy, error bounds, and computational efficiency.
Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.
This work addresses the challenge of statically quantifying rounding errors in floating-point computations. We introduce Λnum, a functional language that—uniquely—integrates sensitivity analysis with graded monads within a linear type system, enabling fully automatic, sound static inference of upper bounds on rounding errors in numerical programs. Λnum natively models IEEE 754 rounding semantics and supports extensions to nondeterministic and stochastic rounding. By rigorously connecting denotational and operational semantics, we establish, for the first time at the type level, soundness guarantees for inferred error bounds. Our prototype implementation demonstrates effectiveness across multiple classical numerical algorithms: it achieves error-bound precision comparable to state-of-the-art tools while significantly accelerating inference speed.
Verifying coverage completeness of input generators in property-based testing remains challenging. Method: This paper proposes a static verification approach based on a “must-style” refinement type system, reformulating conventional “may-produce” type semantics into “must-produce” semantics. It formally defines full coverage for higher-order functions and inductive data types, enabling fully automated verification of generator completeness. Contribution/Results: To our knowledge, this is the first refinement type system provably guaranteeing generation of all inputs satisfying both type and constraint specifications. Experimental evaluation demonstrates substantial improvements in detecting coverage gaps across diverse complex generators, while significantly reducing manual verification effort.
该研究针对SMT求解器处理浮点公式时的性能问题,提出了一种基于语义保持重写规则的变体测试方法,并通过实际测试输入揭示了显著的性能下降现象。
FPScan通过基于抽象解释的静态分析和SMT求解器检测浮点程序中的吸收和灾难性抵消问题。
This work addresses the prevalent overuse of double-precision floating-point numbers in numerical programs by proposing an automated mixed-precision tuning methodology. The approach supports user-defined, non-standard low-precision floating-point formats with customizable exponent and mantissa bit-widths, and integrates numerical validation with systematic search within a unified framework to automatically generate program variants that meet prescribed accuracy constraints. Leveraging the PROMISE tool and containerized parallel benchmarking, the method demonstrates that numerous variables across a range of numerical applications and the Rodinia benchmark suite can be safely downgraded in precision. This reduction yields significant improvements in performance while simultaneously decreasing memory consumption and energy usage, all without compromising numerical accuracy.
This work addresses the challenge of accurately attributing floating-point numerical discrepancies introduced by compilers. To this end, it proposes BMOA, a diagnostic framework that systematically disentangles three dimensions of such deviations: baseline comparison relationships, compiler-induced mechanistic evidence, and precision consequences. Integrating rigorous floating-point semantics, local transformation analysis, cross-compiler comparisons, high-precision validation, and reproducibility testing, BMOA generates auditable and traceable attribution records while preserving ambiguous attributions when evidence is insufficient. Evaluation on six scientific computing kernels targeting the ARM64 platform yielded 1,276 attribution records and 162 mechanistic instances, revealing that the choice of baseline significantly influences diagnostic conclusions and demonstrating that compiler-introduced deviations do not necessarily entail actual precision loss.
Existing Simulink model checkers often produce verification results inconsistent with simulation outcomes due to the absence of bit-precise formal semantics for modeling elements and numerical behaviors, undermining their reliability. This work proposes the first bit-precise conformance testing methodology tailored for Simulink model checkers. By formally specifying the semantics of fundamental blocks, constructing a test suite covering ten block categories, and integrating SMT solving within an automated framework, the approach systematically evaluates behavioral alignment across tools. Experimental results demonstrate that the method effectively uncovers inconsistencies: while the third-party checker SmtMC passes all tests, Simulink Design Verifier exhibits only 94–96% conformance with the simulator, and its agreement with other checkers drops further to 80–90%. The framework also precisely identifies the root causes of these discrepancies.