🤖 AI Summary
This work addresses the challenge of programming hard intellectual property (hard IP) blocks—such as tensor slices—in domain-specific FPGAs, which are typically inaccessible to high-level synthesis (HLS) and require inefficient manual RTL integration. The authors propose a compiler-agnostic approach that leverages HLS black-box mechanisms at the architectural level: hard IP modules are encapsulated as RTL black boxes and modeled as C-level schedulable operators with explicit latency and initiation interval constraints. This enables standard HLS tools, such as AMD Vitis HLS, to directly invoke hard IPs from C/C++ code without compiler modifications or handcrafted co-design. Experiments on a tensor-slice FPGA demonstrate that the proposed method yields designs with lower area-delay product compared to behavioral HLS baselines, while achieving hardware efficiency comparable to hand-written RTL but with substantially improved development productivity.
📝 Abstract
Domain-specific Field Programmable Gate Array (FPGA) architectures increasingly integrate specialized hardblocks, such as Tensor Slices, to accelerate artificial intelligence and machine learning workloads. Despite their efficiency benefits, these architectures remain difficult to program because designers typically rely on manual Register-Transfer Level (RTL) integration to access these hardblocks. This paper presents a compiler-agnostic methodology that enables high-level synthesis (HLS) tools to target custom FPGA hardblocks directly from C/C++ code. Architectural hardblocks are exposed as schedulable C-level operators using an RTL blackbox abstraction with explicit latency and initiation-interval contracts, allowing the HLS scheduler to optimize around specialized hardware without manual RTL orchestration. Unlike traditional uses of HLS blackboxes for external IP integration, our approach treats blackboxes as architectural abstractions, enabling scalable composition of C-level operators that target custom FPGA hardblocks without compiler modification. We evaluate the proposed flow using a Tensor Slice-based FPGA architecture with AMD Vitis HLS and the Verilog-to-Routing (VTR) toolchain. Across multiple matrix sizes, designs generated using the proposed C-Blackbox flow achieve lower area-delay product than behavioral HLS baselines while providing substantially higher productivity-adjusted efficiency than handwritten RTL implementations. These results demonstrate that domain-specific FPGA architectures can be made accessible through HLS while maintaining competitive hardware efficiency.