bram-dsp-aware placement

Designs and implements FPGA placement techniques that detect and create BRAM–DSP macros and co-locate block RAMs and DSPs so placement can use intra-block or direct BRAM‑to‑DSP links instead of global interconnect. Builds placement cost functions, constraints and transformations that exploit those direct paths to reduce critical‑path routing, wirelength and congestion while preserving baseline quality‑of‑results for other designs.

bram-dsp-awareplacement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the performance bottleneck in deep learning accelerators on FPGAs caused by data transfers between Block RAMs (BRAMs) and Digital Signal Processing units (DSPs), which rely heavily on global routing resources, leading to increased wirelength, congestion, and critical-path delay. To mitigate this issue, the authors propose a lightweight architectural enhancement that introduces dedicated direct interconnects between BRAMs and DSPs without altering the overall FPGA architecture or compromising compatibility with existing CAD tools. A complementary placement algorithm is developed to identify and optimize BRAM-DSP macro blocks, substantially reducing reliance on global interconnects. Evaluated on an Agilex-10–like FPGA fabric, the approach achieves up to a 25% improvement in maximum operating frequency (Fmax) and a 49% reduction in wirelength for deep learning layers, while imposing no adverse effects on non-deep-learning benchmarks.

BRAMdata movementDSP

DAS-MP: Enabling High-Quality Macro Placement with Enhanced Dataflow Awareness

May 22, 2025
XZ
Xiaotian Zhao
🏛️ University of Michigan | Shanghai Jiao Tong University | Xi’an Jiaotong-Liverpool University

Existing automated macro placement methods inadequately model data flow—particularly overlooking implicit inter-cluster data dependencies between macros and standard-cell clusters. Method: This work introduces, for the first time, a cross-granularity data-flow modeling framework that explicitly captures and converts such implicit data flows into optimizable placement constraints. It further proposes a congestion-aware macro area-balancing optimization mechanism and a direction-sensitive, data-flow-driven macro flipping algorithm. Results: Evaluated on mainstream benchmarks, the method achieves an average 7.9% reduction in half-perimeter wirelength (HPWL), an 82.5% decrease in congestion overflow, and improvements of 36.97% and 59.44% in worst negative slack (WNS) and total negative slack (TNS), respectively, with negligible runtime overhead (<1.5%). This work presents the first systematic solution for intelligent chip placement that jointly incorporates fine-grained data-flow semantics and physical implementation constraints.

Addressing overlooked standard cell cluster impact on macro placementEnhancing macro placement quality with improved dataflow awarenessOptimizing congestion and wirelength via dataflow-driven fine-tuning steps

This work addresses the limitations of existing FPGA placement tools, which rely on two-dimensional frameworks and struggle to effectively optimize the inter-layer timing and routing characteristics unique to 3D FPGAs. The paper presents the first complete placement flow specifically designed for 3D FPGAs, integrating partition-based initialization, adaptive cost scheduling, fine-grained delay modeling, and a 3D-aware simulated annealing move strategy to jointly optimize layer assignment and timing. Experimental results across four representative 3D architectures demonstrate that the proposed method reduces critical path delay by 2%–6% on average (up to 18%) and decreases total wirelength by 1%–5% on average (up to 10%), significantly improving both timing and routing quality.

3D FPGAdelay modelingplacement

Chip Placement with Diffusion Models

Jul 17, 2024
VL
Vint Lee
🏛️ UC Berkeley

This work addresses the macro-placement optimization problem in digital circuit design by proposing the first diffusion-model-based zero-shot chip placement method. To overcome the poor generalization and low efficiency of conventional reinforcement learning approaches, we design a scalable denoising U-Net architecture integrated with placement-quality-driven conditional guidance and synthetic-data pretraining—enabling cross-circuit zero-shot transfer without task-specific fine-tuning. Evaluated on real-world circuit benchmarks, our method achieves state-of-the-art performance: 12.3% reduction in routing congestion, 8.7% reduction in timing-critical path delay, and significantly superior placement quality compared to both learning-based and heuristic methods. The core contribution is the pioneering application of diffusion models to chip physical design, establishing a new paradigm for high-fidelity, generalizable, and training-free placement generation.

Enabling zero-shot transfer to real circuits using synthetic dataOptimizing macro placement in digital circuit designOvercoming limitations of reinforcement learning in chip placement

Latest Papers

What's happening recently
View more

This work addresses the high latency and scarce interposer resource consumption caused by super long lines (SLLs) in multi-die FPGAs, which often become critical-path bottlenecks. It presents the first logic resynthesis approach that explicitly leverages die-partitioning information during logic synthesis, proposing an interconnect-aware, LUT-level transformation that simplifies local circuit structures to reduce SLL usage. Integrated into a complete FPGA CAD flow encompassing packing and placement, the method achieves up to 24.8% and 27.38% SLL reduction on EPFL benchmarks for 2-die and 3-die configurations, respectively. On MCNC benchmarks, it yields an average 1.65% SLL reduction without degrading placement quality, and significantly lowers inter-die connectivity in Koios designs, thereby enhancing physical design flexibility.

critical pathsinterconnect overheadinterposer resources

This work addresses the underutilization of wide DSP data paths in low-bit quantized neural networks deployed on FPGAs, where existing packing strategies are constrained by fixed bit widths or require substantial external logic. The authors propose a novel dynamic packing technique leveraging the internal pre-adder of DSP blocks, enabling, for the first time, efficient reuse of wide multiplier resources for arbitrary signed or unsigned input bit widths. A customized accelerator architecture tailored for matrix-vector multiplication and convolution is co-designed with this packing scheme. Deeply integrated into the AMD FINN framework, the approach significantly reduces external logic overhead. Evaluated on the UltraNet model, it achieves a 21% reduction in LUT usage and a 36% improvement in frames-per-second per DSP (FPS/DSP) compared to the FINN baseline.

bitwidth packingdatapath underutilizationDSP

This work addresses the significant computational overhead of legality checking during FPGA logic packing, which dominates the CAD flow due to complex logic elements and intricate local interconnect structures. Observing abundant repetitive patterns in the packing process, the authors propose an acceleration method based on pattern memoization, introducing a novel data structure called the “packing signature tree” to efficiently identify and cache legality outcomes of recurring patterns, thereby eliminating redundant computations. The approach is further enhanced by integrating multi-source multi-sink routing-aware legality checks. Evaluated on AMD 7-series and Stratix 10 architectures, the method achieves average packing speedups of 3.7× and 6.9× (up to 13.4× and 29.3×, respectively) and end-to-end VPR flow accelerations of 1.6× and 5.3×, all without degrading solution quality.

CAD runtimeFPGA packinglogic clustering

This work proposes OrderPlace, a novel framework that treats macro placement order as a learnable optimization dimension, overcoming the limitations of traditional static heuristics which often lead to irreversible suboptimal constraints due to early placement decisions. By integrating large language model–guided evolutionary algorithms, code-level policy generation, and a lightweight surrogate evaluator, OrderPlace efficiently explores dynamic and diverse placement sequencing strategies. Evaluated on the ISPD 2005 benchmark suite, the method reduces wirelength by 34.04% and 14.08% compared to WireMask-EA and EGPlace, respectively, demonstrating the superior effectiveness of the discovered placement policies.

chip physical designcombinatorial optimizationmacro placement

This work addresses the inefficiencies of conventional B+ tree search on FPGAs, which suffers from frequent memory accesses, low node reuse, and limited parallelism. To overcome these challenges, the authors propose a batched B+ tree search method tailored for FPGA implementation, employing a layer-wise traversal strategy that processes multiple query keys simultaneously. This approach substantially improves node cache reuse and reduces global memory traffic. A configurable search kernel is developed using high-level synthesis (HLS), enabling flexible tuning of batch size, node width, and tree depth. The design leverages on-chip parallel comparison and node reuse mechanisms on an AMD Alveo U250 FPGA. Experimental results demonstrate that the single-core FPGA implementation achieves a 4.9× speedup over a single-threaded CPU baseline on million-scale B+ trees, while a four-core configuration outperforms a 16-thread CPU by 2.1×.

B+ treebatch searchFPGA

Hot Scholars

EK

Emre Karabulut

North Carolina State University | Microsoft
Hardware SecurityPost-Quantum Cryptography
AA

Aydin Aysu

North Carolina State University
hardware securitypost-quantum cryptographyAI/ML security
AA

Arsalan Ali Malik

North Carolina State University
Fault Injection AttacksComputer ArchitectureEmbedded Systems SecuritySide-Channel Attacks
VL

Vasileios Leon

National Technical University of Athens, School of Electrical & Computer Engineering
HW AccelerationDigital IC DesignEmbedded SystemsSoC/FPGA