Shortest-Path FFT: Optimal SIMD Instruction Scheduling via Graph Search

๐Ÿ“… 2026-04-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF

career value

211K/year
๐Ÿค– AI Summary
This work addresses the challenge of automatically selecting the highest-performing FFT implementation among mathematically equivalent variants on a given hardware platform, which requires modeling context-dependent effects such as cache state transitions between instructions. The authors formulate FFT optimization as a shortest-path problem on a directed acyclic graph and introduce a context-aware graph model in which nodes encode the type of preceding operations, explicitly capturing dependencies like cache warm-upโ€”an aspect overlooked by conventional dynamic programming approaches. Leveraging Dijkstraโ€™s algorithm for graph traversal, combined with empirical SIMD instruction cost measurements and NEON performance profiling, the method achieves 29.8 GFLOPS on an Apple M1 processor, outperforming a pure radix-2 implementation by 5.2ร— and surpassing context-agnostic strategies by 34%. The approach also uncovers highly efficient scheduling sequences, such as R4โ†’R2โ†’R4โ†’R4โ†’Fused-8.

Technology Category

Application Category

๐Ÿ“ Abstract
An $N$-point FFT admits many valid implementations that differ in radix choice, stage ordering, and register-blocking strategy. These alternatives use different SIMD instruction mixes with different latencies, yet produce the same mathematical result. We show that finding the fastest implementation is a shortest-path problem on a directed acyclic graph. We formalize two variants of this graph. In the \emph{context-free} model, nodes represent computation stages and edge weights are independently measured instruction costs. In the \emph{context-aware} model, nodes are expanded to encode the \emph{predecessor edge type}, so that edge weights capture inter-operation correlations such as cache warming -- the cost of operation~B depends on which operation~A preceded it. This addresses a limitation identified but deliberately bypassed by FFTW \citep{FrigoJohnson1998}: that optimal-substructure assumptions break down ``because of the different states of the cache.'' Applied to Apple M1 NEON, the context-free Dijkstra finds an arrangement at 22.1~GFLOPS (74\% of optimal). The context-aware Dijkstra discovers $\text{R4} \to \text{R2} \to \text{R4} \to \text{R4} \to \text{Fused-8}$ at 29.8~GFLOPS -- a $5.2\times$ improvement over pure radix-2 and 34\% faster than the context-free result. This arrangement includes a radix-2 pass \emph{sandwiched between} radix-4 passes, exploiting cache residuals that only exist in context. No context-free search can discover this.
Problem

Research questions and friction points this paper is trying to address.

FFT
SIMD
instruction scheduling
shortest-path
cache effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shortest-Path FFT
SIMD instruction scheduling
context-aware optimization
cache-aware FFT
graph search
๐Ÿ”Ž Similar Papers
No similar papers found.