Score
Designs and implements system-level architectures that integrate multiple FPGA devices into a coordinated computing platform, including partitioning logic across FPGAs, defining low-latency interconnects and communication fabrics, and specifying orchestration between hardware and software stacks. Builds distributed FPGA frameworks and TDAQ-style architectures and develops the device-level drivers, runtime services, configuration and performance/timing analyses needed to operate, monitor, and maintain multi-FPGA deployments.
Heterogeneous FPGAs—integrating DSPs, memory blocks, and domain-specific accelerators—improve PPA but exacerbate design-space exploration complexity due to tight resource coupling. To address this, we propose a lightweight consistency-aware analysis method inspired by the Roofline model, enabling early-stage co-optimization across heterogeneous resources. We introduce three novel consistency metrics to rapidly identify utilization bottlenecks in critical resources such as DSPs and memories. Evaluated on the Stratix 10 FPGA using Koios and VPR benchmarks, our approach performs low-overhead performance profiling. It significantly reduces architectural exploration complexity while preserving modeling efficiency, thereby effectively supporting architecture-application co-optimization for emerging workloads—including machine learning—without sacrificing prediction accuracy or design productivity.
Communication behavior in FPGA data-path designs is difficult to analyze statically, hindering dataflow optimization and parallelism exploitation. Method: This paper proposes the first static communication modeling and analysis framework tailored for hardware data paths. It constructs a pre-RTL communication model based on formal dataflow graphs, integrating static dependency analysis with bandwidth estimation to enable inferable characterization of inter-module data interaction patterns. Contribution/Results: Evaluated across multiple FPGA acceleration benchmarks, the framework achieves an average 18% reduction in routing congestion and a 12% reduction in critical path delay, significantly improving post-synthesis performance and resource utilization. This work overcomes the long-standing limitation of communication non-analyzability in conventional hardware design, establishing both a theoretical foundation and a practical toolset for dataflow-driven architectural optimization.
To address fundamental challenges in FPGA integration within data centers—including low abstraction levels, complex interfaces, and inefficient dynamic partial reconfiguration (DPR)—this paper introduces an open-source shell system for FPGA accelerators. We propose a novel three-layer hierarchical architecture enabling fine-grained, service- and user-logic-aware DPR; provide a unified logical interface; support multithreaded/multi-tenant abstractions; integrate a RoCE v2 network stack; implement an FPGA-GPU DMA engine; enable shared virtual memory; and deliver a high-level programming framework. Experimental evaluation demonstrates a 15–20% reduction in synthesis time and a 10× acceleration in runtime reconfiguration latency. The system successfully deploys real-world applications—including HyperLogLog, AES encryption, and neural network inference—with seamless Python-side invocation. This work significantly enhances FPGA programmability, reusability, and deployment efficiency in heterogeneous computing systems.
To address the challenges of hardware-software co-development and debugging in PCIe-interconnected FPGA systems—including prolonged verification cycles, system instability, and difficulty in error localization—this paper proposes VM-HDL co-simulation, the first virtual machine–HDL co-simulation framework ensuring full-stack consistency. The framework integrates QEMU-based virtualization, SystemC/TLM modeling, a configurable PCIe protocol stack, and a hardware-software time-synchronization mechanism, enabling, for the first time, synchronized debugging and real-time signal observation of OS-level software (including kernel and drivers) alongside RTL hardware. Evaluated on Xilinx Kintex-7 FPGAs running Linux, the framework reduces debugging iteration time by over one order of magnitude, significantly improving development efficiency and system stability. It establishes a new paradigm for high-fidelity, reproducible joint verification of heterogeneous acceleration systems.
This work addresses the challenge that a single FPGA lacks sufficient resources to support full-scale simulation of large-scale multicore RISC-V systems. To overcome this limitation, the authors propose a scalable, multi-FPGA distributed simulation framework that seamlessly maps multicore architectures onto multiple FPGAs through system-level partitioning and an efficient inter-FPGA communication mechanism, enabling cycle-accurate cosimulation without any modifications to the original RTL. Notably, this approach is the first to allow flexible scaling of both core count and FPGA quantity while preserving full system functionality and without requiring RTL reconfiguration. The effectiveness of the framework is demonstrated by successfully simulating a 64-core RISC-V system across eight Alveo U55c FPGAs, achieving a complete Linux boot and validating both functional correctness and scalability of the entire system.
To address low FPGA SoC computational resource utilization and inflexible functional reconfiguration in 5G/6G radio units, this paper proposes a hierarchical, data-driven micro-orchestrator framework supporting event-triggered dynamic partial reconfiguration and hardware resource virtualization. The framework automates the FPGA functional lifecycle management, enabling fine-grained, context-aware, function-level on-demand reconfiguration. Its effectiveness is validated in edge computing applications, particularly computer vision. Compared to conventional static deployment, the proposed approach achieves a measured 42% improvement in hardware resource utilization and reduces reconfiguration latency to under 10 ms, thereby significantly enhancing system real-time responsiveness. This work delivers a lightweight, efficient, and adaptive resource management layer for reconfigurable hardware infrastructures in edge intelligence scenarios.
This work addresses the challenge of programming hard intellectual property (hard IP) blocks—such as tensor slices—in domain-specific FPGAs, which are typically inaccessible to high-level synthesis (HLS) and require inefficient manual RTL integration. The authors propose a compiler-agnostic approach that leverages HLS black-box mechanisms at the architectural level: hard IP modules are encapsulated as RTL black boxes and modeled as C-level schedulable operators with explicit latency and initiation interval constraints. This enables standard HLS tools, such as AMD Vitis HLS, to directly invoke hard IPs from C/C++ code without compiler modifications or handcrafted co-design. Experiments on a tensor-slice FPGA demonstrate that the proposed method yields designs with lower area-delay product compared to behavioral HLS baselines, while achieving hardware efficiency comparable to hand-written RTL but with substantially improved development productivity.
This work addresses the challenges posed by the lengthy and platform-specific nature of FPGA hardware design, which hinders flexible deployment in high-throughput, agile network infrastructures. To overcome this limitation, the authors propose PAF, a framework that employs a pipeline-oriented, parameterized architectural design methodology using the Chisel hardware construction language. By decoupling architectural intent from low-level implementation details, PAF enables fine-grained hardware control while substantially reducing code complexity. The framework facilitates efficient reuse and automated optimization of the same network application across diverse FPGA platforms. Experimental evaluation on an industrial-grade packet classification system demonstrates that PAF achieves performance and resource utilization comparable to hand-optimized designs, while significantly improving development productivity and portability.
Existing FPGA cloud scheduling research lacks standardized, preemptive benchmarking frameworks; evaluations relying on private or synthetic workloads suffer from poor comparability and irreproducibility. Method: We propose FPGA-Bench—the first open-source, scalable, preemptive FPGA benchmark suite—comprising 27 real-world applications across cryptography, AI/ML, and communications. It introduces a novel hardware-level context snapshotting and restoration mechanism, enabling fine-grained resource monitoring and seamless integration with standard scheduler interfaces. Contribution/Results: FPGA-Bench enables repeatable, quantitative evaluation of scheduling fairness, resource efficiency, and preemption overhead in multi-tenant FPGA clouds. By significantly lowering the barrier to preemption strategy validation, it has already facilitated multiple studies on scheduler–OS co-design and advances standardization efforts for FPGA cloud infrastructure evaluation.