π€ AI Summary
This study addresses the loss of precision caused by rounding errors in finite element computations, which remains difficult to analyze a priori. We propose the first automated a posteriori rounding error estimation framework that leverages running error analysis to track numerical and error propagation in real time. Built upon the FEniCS Form Compiler, this method achieves the first automated error estimation within finite element kernels through a C++ backend, custom arithmetic types, and templated kernel generation techniques. Experimental results demonstrate that the framework successfully detects catastrophic cancellation with only a 2β4Γ performance overhead. By effectively supporting mixed-precision design and numerical debugging, this work establishes a new paradigm for ensuring reliability in scientific computing.
π Abstract
Rounding errors in finite element computations can lead to a complete loss of accuracy, stalled convergence, and incorrect results. Moreover, the effects of rounding errors accumulated within automatically generated and compiled kernels are difficult to analyze a priori. We present the first software framework for automated rounding error estimation within finite element kernels. The proposed methodology is based on an a posteriori technique called Running Error Analysis (REA), where a forward error estimate is automatically computed concurrently with the value. An open-source implementation is provided for the FEniCS Form Compiler (FFCx), based on a C++ backend for generating type-generic templated kernels over a custom arithmetic type that tracks both the value and its error estimate.
We demonstrate REA on two examples. First, we use it to detect catastrophic cancellation in the assembly of a Neo-Hooke hyperelastic model in the small deformation regime. A series expansion circumvents the cancellation problem and the computed error estimates show this. Second, we study the assembly of the Laplace operator on a near-degenerate mesh. We demonstrate that our REA implementation typically incurs only 2-4x performance overhead. Applications of this work include robust reduced-precision computations in embedded systems, numerical debugging of new, possibly ill-conditioned or unstable PDE formulations, and guiding the design of mixed-precision kernels.