Score
Designs and implements FPGA microarchitectures and the synthesis/place-and-route toolchain that transforms hardware descriptions into placed-and-routed bitstreams, covering logic synthesis, technology mapping, placement, routing, timing closure, and bitstream generation. Analyzes and optimizes hardware mapping and physical constraints to meet area, timing, power, and resource-utilization targets across synthesis, placement, and routing stages.
This study addresses the challenge of systematically comparing timing behavior of RISC-V processors across heterogeneous technology platforms—specifically, 20 nm FPGAs versus 7 nm FinFET ASICs. We propose a microarchitectural-level, cross-platform timing attribution methodology that integrates static timing analysis (STA), PVT-corner statistical characterization, and pipeline-stage decoupled modeling. Our approach establishes a three-component decomposition framework—logic, routing, and clock—and precisely localizes timing-critical transitions to individual pipeline stages. For the first time, we reveal that FPGA timing is dominated by routing parasitics and topology sensitivity, yielding wide yet scattered timing margins; in contrast, ASIC timing is governed by combinational logic depth and PVT stability, resulting in narrow, concentrated margins. Quantitatively, we identify the EX→MEM stage transition as the common critical path across both platforms. Based on this insight, we formulate predictive, heterogeneity-aware design guidelines for timing convergence.
This work addresses the limitations of existing FPGA placement tools, which rely on two-dimensional frameworks and struggle to effectively optimize the inter-layer timing and routing characteristics unique to 3D FPGAs. The paper presents the first complete placement flow specifically designed for 3D FPGAs, integrating partition-based initialization, adaptive cost scheduling, fine-grained delay modeling, and a 3D-aware simulated annealing move strategy to jointly optimize layer assignment and timing. Experimental results across four representative 3D architectures demonstrate that the proposed method reduces critical path delay by 2%–6% on average (up to 18%) and decreases total wirelength by 1%–5% on average (up to 10%), significantly improving both timing and routing quality.
Existing 3D FPGA research is hindered by fixed prototypes, limited architectural templates, and purely simulation-based evaluation, impeding practical design exploration. This paper presents the first open-source, end-to-end automated framework for 3D FPGA architecture generation and verification—spanning high-level architectural specification, synthesizable RTL generation, and bitstream compilation. Our method introduces customizable vertical interconnect patterns, a novel 3D switch block, and heterogeneous logic-layer architectures, while integrating physical constraint modeling—including through-silicon via (TSV) density and vertical interconnect delay. A closed-loop workflow integrates high-level modeling, RTL synthesis, bitstream generation, physical feasibility validation, and multi-dimensional quantitative assessment to enable efficient architecture-space exploration. Experimental evaluation across five case studies demonstrates significant improvements in average wirelength, critical-path delay, and routing runtime—validating the framework’s effectiveness, scalability, and physical realizability.
To address the lack of explicit temporal semantics and execution-order modeling in functional languages for high-level synthesis (HLS), this paper proposes a novel hardware description methodology integrating the Synchronous Dataflow with Actor Parameters (SDF-AP) model and Template Haskell. It is the first to embed SDF-AP’s production/consumption timing constraints directly into functional specifications, leveraging higher-order function reuse and dataflow patterns to jointly characterize resource allocation and critical-path latency. Built upon the Clash compiler framework, the approach automatically generates VHDL/Verilog code featuring deterministic timing behavior and complete control- and data-path implementations. Experimental evaluation across multiple benchmarks demonstrates stable resource utilization and strict cycle-accurate timing predictability. Compared to Vitis HLS, the method achieves 23–41% average latency reduction and up to 18% lower resource consumption in selected designs.
Existing hardware synthesis approaches decouple implementation selection from scheduling, failing to fully exploit FPGA heterogeneity and yielding suboptimal designs. This paper proposes the first holistic synthesis framework that jointly optimizes implementation selection and scheduling. It employs an e-graph to uniformly model algebraic transformations and hardware implementation decisions, leverages equivalence saturation for efficient exploration of multiple implementation paths, and performs timing-constrained scheduling via a primary mixed-integer linear programming (MILP) formulation augmented by ASAP-based heuristics. Evaluated on Xilinx Kintex UltraScale+ FPGAs, the method achieves an average speedup of 3.01× over Vitis HLS, with up to 5.22× acceleration for complex expressions. This work represents the first systematic breakthrough in hardware synthesis enabling end-to-end co-optimization of implementation and scheduling.
This work addresses the limitations of existing high-level synthesis (HLS) tools in balancing sequential semantics with fine-grained control over pipeline design, which hinders optimization of power, performance, and area (PPA). The paper proposes a novel HLS approach based on visibility control that preserves a sequential programming model while enabling precise manipulation of pipeline structures and hazard-handling mechanisms through a unified visibility abstraction. This framework encompasses strategies such as stall insertion, bypassing, speculative execution, delayed commit, and register renaming. Experimental results on a RISC-V core, histogram computation, and an AES accelerator demonstrate that the generated pipelines significantly outperform those from state-of-the-art sequential-semantics-preserving HLS tools, achieving PPA metrics close to hand-optimized RTL implementations and enabling efficient design space exploration.
This work addresses the limitation of traditional FPGA mapping, which performs dual-output packing only after single-output LUT mapping, thereby overlooking pairing optimization opportunities during cut selection. The authors propose an iterative, dual-output-aware LUT mapping framework that, for the first time, feeds dual-output pairing information forward into the cut selection phase and integrates it into Berkeley ABC. The method alternates between cut selection and constrained dual-output matching by generating candidate pairs via sparse support indexing, scoring matches heuristically, adjusting cut costs with compatibility awareness, and validating architectural legality and timing based on physical input unions. Evaluated on EPFL benchmarks, the approach reduces LUT area by 34.96% on average compared to the original ABC, achieves a 15.8× speedup over the previous best method, and further lowers circuit depth by approximately 5% while reducing area by an additional 1%.
A “dual-language gap” persists between algorithm development in high-level languages and hardware implementation in low-level HDLs. Method: This paper introduces the first MLIR-based high-level synthesis (HLS) toolchain natively supporting Julia—compiling Julia kernels directly to vendor-agnostic, synthesizable SystemVerilog RTL without language extensions or manual annotations, while natively integrating AXI4-Stream protocol support. It innovatively enables hybrid static-dynamic scheduling to balance expressiveness and controllability. Contribution/Results: The generated RTL operates stably at 100 MHz on FPGA. On signal processing and mathematical benchmarks, throughput reaches 59.71%–82.6% of leading C/C++ HLS tools. This significantly improves end-to-end development efficiency and hardware portability—from algorithm specification to synthesized RTL—while preserving Julia’s composability and productivity.