๐ค AI Summary
This work addresses the challenge of automatically selecting the highest-performing FFT implementation among mathematically equivalent variants on a given hardware platform, which requires modeling context-dependent effects such as cache state transitions between instructions. The authors formulate FFT optimization as a shortest-path problem on a directed acyclic graph and introduce a context-aware graph model in which nodes encode the type of preceding operations, explicitly capturing dependencies like cache warm-upโan aspect overlooked by conventional dynamic programming approaches. Leveraging Dijkstraโs algorithm for graph traversal, combined with empirical SIMD instruction cost measurements and NEON performance profiling, the method achieves 29.8 GFLOPS on an Apple M1 processor, outperforming a pure radix-2 implementation by 5.2ร and surpassing context-agnostic strategies by 34%. The approach also uncovers highly efficient scheduling sequences, such as R4โR2โR4โR4โFused-8.
๐ Abstract
An $N$-point FFT admits many valid implementations that differ in radix choice, stage ordering, and register-blocking strategy. These alternatives use different SIMD instruction mixes with different latencies, yet produce the same mathematical result. We show that finding the fastest implementation is a shortest-path problem on a directed acyclic graph.
We formalize two variants of this graph. In the \emph{context-free} model, nodes represent computation stages and edge weights are independently measured instruction costs. In the \emph{context-aware} model, nodes are expanded to encode the \emph{predecessor edge type}, so that edge weights capture inter-operation correlations such as cache warming -- the cost of operation~B depends on which operation~A preceded it. This addresses a limitation identified but deliberately bypassed by FFTW \citep{FrigoJohnson1998}: that optimal-substructure assumptions break down ``because of the different states of the cache.''
Applied to Apple M1 NEON, the context-free Dijkstra finds an arrangement at 22.1~GFLOPS (74\% of optimal). The context-aware Dijkstra discovers $\text{R4} \to \text{R2} \to \text{R4} \to \text{R4} \to \text{Fused-8}$ at 29.8~GFLOPS -- a $5.2\times$ improvement over pure radix-2 and 34\% faster than the context-free result. This arrangement includes a radix-2 pass \emph{sandwiched between} radix-4 passes, exploiting cache residuals that only exist in context. No context-free search can discover this.