Score
Designs, implements, and evaluates algorithmic methods, checks, and tests that keep numerical computations well behaved under finite-precision arithmetic; this includes diagnosing conditioning and singularities, performing rounding-error analysis, evaluating error growth across kernels, and running representative stability benchmarks. It also develops fixes and tunings such as regularization of ill-conditioned problems and matrix inversions, enforcement of Hermitian positive semidefinite structure, stabilization of autodiff gradients and iterative methods, and selection of safe precisions and stability-driven precision assessments.
This work addresses the problem of deriving provably tight floating-point rounding error bounds for numerical programs featuring conditional branches, no loops, and mixed-precision arithmetic. Methodologically, it unifies the modeling of conditional control flow and precision heterogeneity via two novel quantitative metrics—“instability jumps” and “window width”—and integrates interval arithmetic, abstract interpretation, and precision-aware semantic modeling, augmented with abstraction-guided global optimization. Its key contribution is the first formal framework enabling joint, compositional analysis of conditional branching and mixed precision, achieving both high bound tightness and practical analysis efficiency. Experimental evaluation on standard benchmarks demonstrates significantly tighter error bounds compared to prior approaches. Furthermore, the framework successfully guides precision configuration—e.g., step size and search direction—in the conjugate gradient method, empirically validating its utility in supporting design-time trade-offs among accuracy, error bounds, and computational efficiency.
This work proposes a systematic framework to enhance the computational efficiency and numerical stability of evaluating high-degree matrix polynomials. Specifically, for polynomial degrees eight and higher, the method generates and validates stable coefficient sets that reduce the number of required matrix multiplications by one compared to the classical Paterson–Stockmeyer scheme. To address instability issues in the original formulation, the authors introduce structural variants and design a reliability metric to assess the expected numerical accuracy of candidate coefficient sets. Nonlinear polynomial systems are solved using variable-precision arithmetic (VPA), and an in-house tool, MatrixPolEval1, enables efficient screening and validation. Applied to matrix exponentials and geometric series, the approach achieves a saving of one matrix multiplication while maintaining comparable numerical accuracy.
Existing approaches lack the capability to perform automated backward error analysis for numerical programs, making it difficult to verify their backward stability. This work proposes a formal framework that generalizes the definition of backward stability, introduces the category Shel to model stable numerical computations, and develops the tool eggshel to automatically synthesize error bounds. The framework incorporates a novel, composable, and flexible notion of stability, integrating category theory, formal verification, and symbolic reasoning to automatically search for stability proofs within subcategories of Shel. Notably, eggshel is the first tool capable of automating the analysis of programs with variable reuse, successfully generating backward error bounds for several numerical programs previously beyond the reach of existing methods, while providing formal correctness guarantees.
This work addresses the prevalent overuse of double-precision floating-point numbers in numerical programs by proposing an automated mixed-precision tuning methodology. The approach supports user-defined, non-standard low-precision floating-point formats with customizable exponent and mantissa bit-widths, and integrates numerical validation with systematic search within a unified framework to automatically generate program variants that meet prescribed accuracy constraints. Leveraging the PROMISE tool and containerized parallel benchmarking, the method demonstrates that numerous variables across a range of numerical applications and the Rodinia benchmark suite can be safely downgraded in precision. This reduction yields significant improvements in performance while simultaneously decreasing memory consumption and energy usage, all without compromising numerical accuracy.
This work addresses the challenge of statically quantifying rounding errors in floating-point computations. We introduce Λnum, a functional language that—uniquely—integrates sensitivity analysis with graded monads within a linear type system, enabling fully automatic, sound static inference of upper bounds on rounding errors in numerical programs. Λnum natively models IEEE 754 rounding semantics and supports extensions to nondeterministic and stochastic rounding. By rigorously connecting denotational and operational semantics, we establish, for the first time at the type level, soundness guarantees for inferred error bounds. Our prototype implementation demonstrates effectiveness across multiple classical numerical algorithms: it achieves error-bound precision comparable to state-of-the-art tools while significantly accelerating inference speed.
This study addresses the challenge of numerical instability in scientific software caused by floating-point precision errors, particularly in safety-critical contexts where traditional methods struggle with complex expressions. It presents the first systematic evaluation of large language models (LLMs) for improving numerical stability by detecting and rewriting unstable arithmetic expressions. The experiments encompass 2,470 expressions featuring nested conditionals, high-precision literals, and multi-variable arithmetic, evaluated across six prominent LLMs. Results demonstrate that LLMs outperform baseline methods in 65.4% of cases and successfully stabilize 97.9% of the 431 instances where baselines completely fail. Nevertheless, limitations persist in handling control flow constructs and high-precision literals, highlighting areas for future improvement.
This study investigates whether stochastic rounding (SR) retains its regularizing effect in matrices with constant aspect ratios and examines its impact on the singular value spectrum. By integrating singular value analysis, a stochastic rounding quantization model, and spectral theory, the work demonstrates for the first time that SR not only enhances the smallest singular value but also collectively elevates multiple singular values in the tail of the spectrum. This finding reveals that the regularizing influence of SR extends beyond extreme aspect ratio regimes. The results establish SR as a universal spectral regularization mechanism, thereby broadening its theoretical foundation and application potential in numerical computation and low-precision machine learning.
This work addresses the widespread lack of correctly rounded results in high-performance vector math libraries, which undermines bit-level reproducibility across platforms. The authors propose a unified framework that integrates SIMD parallelism with correctly rounded algorithms to efficiently implement multiple single-precision, single-input mathematical functions on CPUs and, for the first time, extend this approach to GPUs. They also provide a prototype implementation for double-precision functions. This research lays the foundation for the first cross-platform vector math library supporting correct rounding, with a planned public release by mid-2026, significantly advancing reproducibility and precision guarantees in numerical computing.
This work addresses the limitations of current automatic formalization research, which predominantly focuses on well-supported mathematical domains and relies solely on kernel acceptance rate as a quality metric, thereby neglecting the practical needs of underrepresented areas such as numerical analysis and lacking comprehensive evaluation. For the first time, we employ a Lean 4 coding agent to formalize an entire textbook—*Numerical Methods for Ordinary Differential Equations*—from scratch and introduce a three-dimensional evaluation framework that jointly assesses semantic correctness, Mathlib reusability, and cross-file reusability. Through LLM-as-judge, semantic validation, and dependency analysis, we uncover pervasive issues in existing systems, including incomplete statements and weakened assumptions, demonstrating that kernel acceptance rate substantially overestimates formalization quality. Our approach establishes a reproducible, multidimensional auditing paradigm for trustworthy automated formalization.