Score
Designs, analyzes, and optimizes FPGA designs to meet timing requirements by identifying and shortening critical paths, applying pipelining and retiming, and iterating synthesis constraints, placement, and floorplanning while trading off area and timing to achieve timing closure.
This study addresses the fundamental differences and challenges in timing closure between FPGAs and ASICs. To tackle timing behavior divergence arising from their architectural heterogeneity, we propose the first cross-platform unified timing analysis framework, integrating static timing analysis, process-node-aware comparison, and silicon measurement validation. Using a comparative case study of Xilinx Kintex UltraScale+ FPGAs and 7 nm ASICs, we quantitatively characterize their timing performance boundaries: the ASIC achieves 45 ps setup and 35 ps hold times, while the FPGA attains 180 ps setup and 120 ps hold times—demonstrating substantial improvements in high-precision timing control for modern high-end FPGAs. The framework systematically elucidates the architecture–timing mapping relationship and enables performance- and programmability-aware co-design decisions in heterogeneous computing platforms. It provides both theoretical foundations and practical guidelines for high-performance reconfigurable systems.
This study addresses the challenge of systematically comparing timing behavior of RISC-V processors across heterogeneous technology platforms—specifically, 20 nm FPGAs versus 7 nm FinFET ASICs. We propose a microarchitectural-level, cross-platform timing attribution methodology that integrates static timing analysis (STA), PVT-corner statistical characterization, and pipeline-stage decoupled modeling. Our approach establishes a three-component decomposition framework—logic, routing, and clock—and precisely localizes timing-critical transitions to individual pipeline stages. For the first time, we reveal that FPGA timing is dominated by routing parasitics and topology sensitivity, yielding wide yet scattered timing margins; in contrast, ASIC timing is governed by combinational logic depth and PVT stability, resulting in narrow, concentrated margins. Quantitatively, we identify the EX→MEM stage transition as the common critical path across both platforms. Based on this insight, we formulate predictive, heterogeneity-aware design guidelines for timing convergence.
In HLS design, significant discrepancies between tool-reported and FPGA-measured clock cycles lead to erroneous bottleneck identification and suboptimal design decisions; existing on-chip analysis methods rely on manual RTL inspection and signal monitoring, resulting in low efficiency. This paper proposes the first fully automated, in-FPGA performance profiling framework tailored for HLS. It leverages pragma-driven RTL generation, incremental synthesis, and cycle-accurate counter insertion to enable non-intrusive, timing-isolated cycle measurement at both single-instruction and full-function granularity, while supporting automated design-space exploration. Evaluated on 28 large-scale designs, the framework achieves 100% cycle capture accuracy with only 16.98% LUT and 43.15% FF overhead, and zero BRAM utilization.
High-level synthesis (HLS) for FPGAs faces challenges in timing closure and architecture-specific optimization, which heavily rely on manual pragma insertion and lack automated, intelligent support. Method: This paper proposes TimelyHLS—a novel framework that integrates large language models (LLMs) with retrieval-augmented generation (RAG) into the HLS flow. It constructs a timing-aware, structured FPGA knowledge base and leverages synthesis log feedback alongside closed-loop evaluation using commercial toolchains to enable automatic, iterative inference and refinement of architecture-specific pragmas. Contribution/Results: Evaluated across ten FPGA platforms, TimelyHLS achieves up to 3.85× speedup for matrix multiplication and 57% register reduction for Viterbi decoding—while guaranteeing functional correctness and timing convergence—significantly reducing human tuning effort.
Existing hardware synthesis approaches decouple implementation selection from scheduling, failing to fully exploit FPGA heterogeneity and yielding suboptimal designs. This paper proposes the first holistic synthesis framework that jointly optimizes implementation selection and scheduling. It employs an e-graph to uniformly model algebraic transformations and hardware implementation decisions, leverages equivalence saturation for efficient exploration of multiple implementation paths, and performs timing-constrained scheduling via a primary mixed-integer linear programming (MILP) formulation augmented by ASAP-based heuristics. Evaluated on Xilinx Kintex UltraScale+ FPGAs, the method achieves an average speedup of 3.01× over Vitis HLS, with up to 5.22× acceleration for complex expressions. This work represents the first systematic breakthrough in hardware synthesis enabling end-to-end co-optimization of implementation and scheduling.
This work addresses the limitations of existing hardware parser designs, which suffer from excessive complexity, poor reusability, and inadequate support for sophisticated matching and diverse deployment scenarios. To overcome these challenges, the authors propose an open-source tool that enhances pattern-matching capabilities through customizable symbolic tokens—enabling range validation, negation, and comparisons with external ports—and introduces a Parser Intermediate Representation (PIR) to decouple frontend protocol specification from backend implementation. The frontend allows flexible protocol description, while the backend automatically generates FPGA-optimized SystemVerilog code supporting arbitrary bit-width state machines, byte alignment, and cross-cycle field stitching. Experimental results on an Ethernet parser demonstrate up to a 226% increase in operating frequency and a 97% reduction in logic resource usage; furthermore, the hierarchical design achieves up to 8× greater resource efficiency compared to monolithic architectures.
Traditional simulation-based dynamic timing analysis struggles to balance accuracy and efficiency, and existing gate delay models lack sufficient expressiveness to enable precise, exhaustive path delay analysis for digital circuits. This work proposes a symbolic execution framework integrated with an analytical gate delay model that automatically generates symbolic delay expressions for all paths under a given input transition ordering. For the first time, it incorporates an analytical delay model accounting for both drafting effects and multi-input switching into symbolic execution. By employing a path-sensitive, goal-directed inference mechanism together with symbolic pruning strategies, the approach significantly enhances the completeness and precision of timing analysis while effectively mitigating the combinatorial explosion problem.
This work presents the first systematic evaluation of large language models (LLMs) for generating latency-sensitive financial FPGA hardware, addressing the need for rapid iteration amid frequent protocol and regulatory changes. The authors introduce FinHardBench, a benchmark comprising 33 tasks, and conduct three types of experiments emulating real-world development workflows: module generation, system-level design space exploration (DSE) of a six-stage trading pipeline, and adaptation to specification changes. Across over 1,530 experiments, six LLMs achieved functional correctness rates of 19–61%, though some exhibited timing performance degradation up to 13.7×. Notably, the best-performing LLM consistently identified the globally optimal configuration in all five system-level DSE trials, significantly outperforming baseline methods such as random search, simulated annealing, and Bayesian optimization. These results reveal an inconsistency between LLMs’ code generation and architectural optimization capabilities, with task difficulty more closely tied to the availability of patterns in training data than to abstraction level.
This work addresses the limitations of existing high-level synthesis (HLS) tools in balancing sequential semantics with fine-grained control over pipeline design, which hinders optimization of power, performance, and area (PPA). The paper proposes a novel HLS approach based on visibility control that preserves a sequential programming model while enabling precise manipulation of pipeline structures and hazard-handling mechanisms through a unified visibility abstraction. This framework encompasses strategies such as stall insertion, bypassing, speculative execution, delayed commit, and register renaming. Experimental results on a RISC-V core, histogram computation, and an AES accelerator demonstrate that the generated pipelines significantly outperform those from state-of-the-art sequential-semantics-preserving HLS tools, achieving PPA metrics close to hand-optimized RTL implementations and enabling efficient design space exploration.
This work addresses the limitation of traditional FPGA mapping, which performs dual-output packing only after single-output LUT mapping, thereby overlooking pairing optimization opportunities during cut selection. The authors propose an iterative, dual-output-aware LUT mapping framework that, for the first time, feeds dual-output pairing information forward into the cut selection phase and integrates it into Berkeley ABC. The method alternates between cut selection and constrained dual-output matching by generating candidate pairs via sparse support indexing, scoring matches heuristically, adjusting cut costs with compatibility awareness, and validating architectural legality and timing based on physical input unions. Evaluated on EPFL benchmarks, the approach reduces LUT area by 34.96% on average compared to the original ABC, achieves a 15.8× speedup over the previous best method, and further lowers circuit depth by approximately 5% while reducing area by an additional 1%.