🤖 AI Summary
To address the high-throughput, numerical stability, and computational efficiency requirements for matrix exponential evaluation in generative AI, this paper proposes a novel adaptive Taylor series evaluation algorithm. Unlike classical approaches—such as Paterson–Stockmeyer or Padé approximants combined with scaling-and-squaring—the method jointly optimizes the Taylor expansion order and scaling factor to minimize floating-point operations under a prescribed relative error tolerance (<1e−12), while integrating dynamic error control within a scaling-evaluation-squaring framework. Theoretical analysis guarantees numerical stability and achieves a Pareto improvement in both accuracy and asymptotic complexity. Empirically, on large-scale generative tasks—including diffusion models and continuous-time flow matching—the algorithm achieves an average 2.3× speedup over state-of-the-art methods. The implementation is open-sourced and validated across mainstream GPU platforms.
📝 Abstract
The matrix exponential is a fundamental operator in scientific computing and system simulation, with applications ranging from control theory and quantum mechanics to modern generative machine learning. While Padé approximants combined with scaling and squaring have long served as the standard, recent Taylor-based methods, which utilize polynomial evaluation schemes that surpass the classical Paterson--Stockmeyer technique, offer superior accuracy and reduced computational complexity. This paper presents an optimized Taylor-based algorithm for the matrix exponential, specifically designed for the high-throughput requirements of generative AI flows. We provide a rigorous error analysis and develop a dynamic selection strategy for the Taylor order and scaling factor to minimize computational effort under a prescribed error tolerance. Extensive numerical experiments demonstrate that our approach provides significant acceleration and maintains high numerical stability compared to existing state-of-the-art implementations. These results establish the proposed method as a highly efficient tool for large-scale generative modeling.