🤖 AI Summary
This study addresses the lack of fine-grained bit-width control and rapid performance evaluation in hardware deployment for sub-microsecond dataflow kernels by proposing Alkaid, an open-source compiler. The method constructs an arithmetic-centric intermediate representation (ALIR) to preserve fixed-point semantics and heterogeneous bit-widths, while introducing a white-box analytical model that enables synthesis-free, rapid design space exploration. Alkaid translates low-latency kernels into RTL/HLS code within seconds, facilitating efficient hardware mapping for both ML pre/post-processing and non-ML logic. Compared with state-of-the-art pipelines, it reduces LUT resource utilization by 20–30% and decreases latency by up to 40% in certain scenarios, successfully integrating non-ML components such as sorting networks.
📝 Abstract
Ultra-low-latency machine learning and data processing pipelines operating on the sub-microsecond level often contain static dataflow kernels that require fine-grained bitwidth control, arithmetic optimization, and fast hardware performance estimation. This works introduce Alkaid, a free and open source domain specific compiler that translates sub-microsecond latency dataflow kernels into platform agnostic Register Transfer Level (RTL) or High-Level Synthesis (HLS)-ready code in seconds. Alkaid targets ultra-low-latency ML pipelines, especially those addressed by existing ML-to-hardware flows such as \texttt{hls4ml}, while also supporting the surrounding pre-processing, post-processing, and control logic. At its core, Alkaid uses an arithmetic centric intermediate representation, Alkaid Low-level IR (ALIR), to preserve fixed-point semantics, heterogeneous bitwidths, and operation level structure for optimization and code generation. Alkaid further provides comprehensive hardware-aware optimization passes and an interpretable white-box analytical performance model for rapid design space exploration without synthesis in the loop. Across representative ML workloads, including MLPs, GNNs, and transformers, Alkaid achieves 20-30\% lower LUT usage than the best prior flows for bit-exact equivalent designs, with up to 40\% lower latency in some cases. Beyond typical neural networks, Alkaid achieves comparable performance to the state-of-the-art method for synthesizing boosted decision trees, while also enabling the implementation and integration of non-ML kernels such as sorting networks and histograms.