fpga synthesis and place-and-route

Designs and implements FPGA microarchitectures and the synthesis/place-and-route toolchain that transforms hardware descriptions into placed-and-routed bitstreams, covering logic synthesis, technology mapping, placement, routing, timing closure, and bitstream generation. Analyzes and optimizes hardware mapping and physical constraints to meet area, timing, power, and resource-utilization targets across synthesis, placement, and routing stages.

fpgasynthesisandplace-and-route

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Pipeline Stage Resolved Timing Characterization of FPGA and ASIC Implementations of a RISC V Processor

Dec 15, 2025
MD
Mostafa Darvishi
🏛️ École de technologie supérieure

This study addresses the challenge of systematically comparing timing behavior of RISC-V processors across heterogeneous technology platforms—specifically, 20 nm FPGAs versus 7 nm FinFET ASICs. We propose a microarchitectural-level, cross-platform timing attribution methodology that integrates static timing analysis (STA), PVT-corner statistical characterization, and pipeline-stage decoupled modeling. Our approach establishes a three-component decomposition framework—logic, routing, and clock—and precisely localizes timing-critical transitions to individual pipeline stages. For the first time, we reveal that FPGA timing is dominated by routing parasitics and topology sensitivity, yielding wide yet scattered timing margins; in contrast, ASIC timing is governed by combinational logic depth and PVT stability, resulting in narrow, concentrated margins. Quantitatively, we identify the EX→MEM stage transition as the common critical path across both platforms. Based on this insight, we formulate predictive, heterogeneity-aware design guidelines for timing convergence.

Characterizes timing of RISC-V processor on FPGA and ASICCompares timing mechanisms across different implementation technologiesIdentifies platform-specific bottlenecks for predictable timing closure

This work addresses the limitations of existing FPGA placement tools, which rely on two-dimensional frameworks and struggle to effectively optimize the inter-layer timing and routing characteristics unique to 3D FPGAs. The paper presents the first complete placement flow specifically designed for 3D FPGAs, integrating partition-based initialization, adaptive cost scheduling, fine-grained delay modeling, and a 3D-aware simulated annealing move strategy to jointly optimize layer assignment and timing. Experimental results across four representative 3D architectures demonstrate that the proposed method reduces critical path delay by 2%–6% on average (up to 18%) and decreases total wirelength by 1%–5% on average (up to 10%), significantly improving both timing and routing quality.

3D FPGAdelay modelingplacement

LaZagna: An Open-Source Framework for Flexible 3D FPGA Architectural Exploration

May 08, 2025
IY
Ismael Youssef
🏛️ Georgia Institute of Technology

Existing 3D FPGA research is hindered by fixed prototypes, limited architectural templates, and purely simulation-based evaluation, impeding practical design exploration. This paper presents the first open-source, end-to-end automated framework for 3D FPGA architecture generation and verification—spanning high-level architectural specification, synthesizable RTL generation, and bitstream compilation. Our method introduces customizable vertical interconnect patterns, a novel 3D switch block, and heterogeneous logic-layer architectures, while integrating physical constraint modeling—including through-silicon via (TSV) density and vertical interconnect delay. A closed-loop workflow integrates high-level modeling, RTL synthesis, bitstream generation, physical feasibility validation, and multi-dimensional quantitative assessment to enable efficient architecture-space exploration. Experimental evaluation across five case studies demonstrates significant improvements in average wirelength, critical-path delay, and routing runtime—validating the framework’s effectiveness, scalability, and physical realizability.

Fixed prototypes and simulation-only evaluations in 3D FPGA studiesLack of open-source frameworks for 3D FPGA architecture generationLimited application of 3D IC technology to FPGAs

High-Level Synthesis of Digital Circuits from Template Haskell and SDF-AP

Apr 10, 2025
HF
H. Folmer
🏛️ University of Twente | Saxion Hogeschool

To address the lack of explicit temporal semantics and execution-order modeling in functional languages for high-level synthesis (HLS), this paper proposes a novel hardware description methodology integrating the Synchronous Dataflow with Actor Parameters (SDF-AP) model and Template Haskell. It is the first to embed SDF-AP’s production/consumption timing constraints directly into functional specifications, leveraging higher-order function reuse and dataflow patterns to jointly characterize resource allocation and critical-path latency. Built upon the Clash compiler framework, the approach automatically generates VHDL/Verilog code featuring deterministic timing behavior and complete control- and data-path implementations. Experimental evaluation across multiple benchmarks demonstrates stable resource utilization and strict cycle-accurate timing predictability. Compared to Vitis HLS, the method achieves 23–41% average latency reduction and up to 18% lower resource consumption in selected designs.

Adding time and execution order to functional HLS descriptionsImproving latency and resource consumption in HLS toolsSynthesizing parallel hardware using SDF-AP patterns

Latest Papers

What's happening recently
View more

Existing hardware synthesis approaches decouple implementation selection from scheduling, failing to fully exploit FPGA heterogeneity and yielding suboptimal designs. This paper proposes the first holistic synthesis framework that jointly optimizes implementation selection and scheduling. It employs an e-graph to uniformly model algebraic transformations and hardware implementation decisions, leverages equivalence saturation for efficient exploration of multiple implementation paths, and performs timing-constrained scheduling via a primary mixed-integer linear programming (MILP) formulation augmented by ASAP-based heuristics. Evaluated on Xilinx Kintex UltraScale+ FPGAs, the method achieves an average speedup of 3.01× over Vitis HLS, with up to 5.22× acceleration for complex expressions. This work represents the first systematic breakthrough in hardware synthesis enabling end-to-end co-optimization of implementation and scheduling.

Exploits FPGA heterogeneous architectures through unified e-graph representationJointly optimizes implementation selection and scheduling for hardware synthesisOvercomes suboptimal designs from separated optimization in current HLS tools

This work addresses the limitations of existing high-level synthesis (HLS) tools in balancing sequential semantics with fine-grained control over pipeline design, which hinders optimization of power, performance, and area (PPA). The paper proposes a novel HLS approach based on visibility control that preserves a sequential programming model while enabling precise manipulation of pipeline structures and hazard-handling mechanisms through a unified visibility abstraction. This framework encompasses strategies such as stall insertion, bypassing, speculative execution, delayed commit, and register renaming. Experimental results on a RISC-V core, histogram computation, and an AES accelerator demonstrate that the generated pipelines significantly outperform those from state-of-the-art sequential-semantics-preserving HLS tools, achieving PPA metrics close to hand-optimized RTL implementations and enabling efficient design space exploration.

Hazard ResolutionHigh-Level SynthesisPipelining

This work addresses the limitation of traditional FPGA mapping, which performs dual-output packing only after single-output LUT mapping, thereby overlooking pairing optimization opportunities during cut selection. The authors propose an iterative, dual-output-aware LUT mapping framework that, for the first time, feeds dual-output pairing information forward into the cut selection phase and integrates it into Berkeley ABC. The method alternates between cut selection and constrained dual-output matching by generating candidate pairs via sparse support indexing, scoring matches heuristically, adjusting cut costs with compatibility awareness, and validating architectural legality and timing based on physical input unions. Evaluated on EPFL benchmarks, the approach reduces LUT area by 34.96% on average compared to the original ABC, achieves a 15.8× speedup over the previous best method, and further lowers circuit depth by approximately 5% while reducing area by an additional 1%.

cut selectiondual-output packingFPGA architecture

A High-level Synthesis Toolchain for the Julia Language

Dec 17, 2025
BS
Benedict Short
🏛️ Imperial College London

A “dual-language gap” persists between algorithm development in high-level languages and hardware implementation in low-level HDLs. Method: This paper introduces the first MLIR-based high-level synthesis (HLS) toolchain natively supporting Julia—compiling Julia kernels directly to vendor-agnostic, synthesizable SystemVerilog RTL without language extensions or manual annotations, while natively integrating AXI4-Stream protocol support. It innovatively enables hybrid static-dynamic scheduling to balance expressiveness and controllability. Contribution/Results: The generated RTL operates stably at 100 MHz on FPGA. On signal processing and mathematical benchmarks, throughput reaches 59.71%–82.6% of leading C/C++ HLS tools. This significantly improves end-to-end development efficiency and hardware portability—from algorithm specification to synthesized RTL—while preserving Julia’s composability and productivity.

Addresses the two-language problem in FPGA accelerator developmentAutomates Julia-to-SystemVerilog compilation for hardware synthesisEnables domain experts to deploy Julia kernels on FPGAs directly

Hot Scholars

WL

Wayne Luk

Professor of Computer Engineering, Imperial College London
Hardware and ArchitectutreReconfigurable ComputingDesign Automation
AB

Andrew Boutros

University of Waterloo
FPGAsReconfigurable ComputingComputer-Aided Design
GS

Gregor Schiele

Professor of Computer Science (Embedded Systems), University Duisburg-Essen, Germany
embedded AIIoTembedded softwareadaptive SW and reconfigurable HW
KJ

Kamil Jeziorek

AGH University of Krakow
Event CamerasGraph Neural NetworksObject DetectionComputer Vision
MS

Muhammad Shafique

Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart