🤖 AI Summary
FDTD-based electromagnetic simulations suffer from poor portability, high development overhead, and performance bottlenecks on modern hardware. To address these challenges, this paper introduces the first MLIR/LLVM-based domain-specific compiler for FDTD. It models the 3D FDTD kernel as semantically explicit 3D tensor operations and proposes a novel high-order tensor abstraction with an automated optimization framework supporting loop tiling, fusion, and vectorization. The compiler features hardware-aware, end-to-end code generation across heterogeneous platforms (x86 and ARM). Experimental evaluation demonstrates up to 10× speedup over NumPy baselines across multiple architectures. By eliminating manual tuning, it overcomes performance fragmentation and non-portability inherent in conventional approaches, thereby significantly improving simulation efficiency, scalability, and deployment flexibility.
📝 Abstract
The Finite Difference Time Domain (FDTD) method is a widely used numerical technique for solving Maxwell's equations, particularly in computational electromagnetics and photonics. It enables accurate modeling of wave propagation in complex media and structures but comes with significant computational challenges. Traditional FDTD implementations rely on handwritten, platform-specific code that optimizes certain kernels while underperforming in others. The lack of portability increases development overhead and creates performance bottlenecks, limiting scalability across modern hardware architectures. To address these challenges, we introduce an end-to-end domain-specific compiler based on the MLIR/LLVM infrastructure for FDTD simulations. Our approach generates efficient and portable code optimized for diverse hardware platforms.We implement the three-dimensional FDTD kernel as operations on a 3D tensor abstraction with explicit computational semantics. High-level optimizations such as loop tiling, fusion, and vectorization are automatically applied by the compiler. We evaluate our customized code generation pipeline on Intel, AMD, and ARM platforms, achieving up to $10 imes$ speedup over baseline Python implementation using NumPy.