🤖 AI Summary
This study addresses throughput bottlenecks in DSP and machine learning workloads caused by cyclic dependencies and multi-cycle computations. To overcome the limitations of conventional decoupled scheduling, this work proposes a recursion-aware temporal mapping method for coarse-grained reconfigurable arrays (CGRAs). The approach optimizes dataflow graph representations via dominant constraints to enable cross-iteration overlapped scheduling. Furthermore, it introduces a unified iteration offset strategy that reduces both initiation intervals and configuration overhead. Experimental results demonstrate that, compared with existing baselines, the proposed method achieves a 2.18× throughput improvement, shortens the initiation interval by 46%, and accelerates mapping convergence by an average factor of 5.07×.
📝 Abstract
Throughput in DSP and machine learning workloads is often limited by two temporal structures, i.e., loop-carried recurrences and long-latency, multi-cycle compute nodes. On spatio-temporal coarse-grained reconfigurable arrays (CGRAs), both bottlenecks can be addressed by overlapping iterations across the multi-context modulo configurations. Yet, existing CGRA mappers schedule a fixed dataflow graph (DFG) that treats recurrence-aware scheduling and operator-level pipelining separately, limiting inter-iteration overlap and inflating routing pressure. To tackle this, we present CONFERM, a recurrence-aware temporal mapper that uses the dominant temporal con-straint to guide the DFG representation and expose opportunities for loop-carried pipelining. CONFERM identifies and prioritizes bottleneck regions during scheduling. The regular loop-carried offsets across interleaved iterations allow the emitted control sequence to repeat at a shorter cadence than the original initiation interval, thus delivering higher throughput with lower CGRA configuration overhead. Across ten benchmark kernels, CONFERM improves throughput by 2.18x over state-of-the-art mappers. Its uniform iteration offsets shorten the emitted initiation interval by 46%. CONFERM's mapper pass also converges faster by 5.07x on average with the same heuristic mapper backend.