Score
Designs and implements the physical implementation of application-specific integrated circuits (ASICs), producing floorplans, cell and block layouts, clock-tree synthesis, and place-and-route results up to GDSII/tape-out. Performs area, power, and timing estimation and optimization, drives signoff and prototyping/fabrication flows, and analyzes hardware-efficiency metrics such as area, power, and timing.
To address the high interconnect overhead and low core density of domain-specific instruction-set processors (DSIPs) for machine learning under Ångström-scale fabrication nodes, this work proposes a physically efficient DSIP architecture that tightly integrates customized near-memory storage structures with compact SIMD compute units, substantially reducing routing complexity. Comprehensive evaluation across five configurations—using the IMEC A10 nanosheet PDK—demonstrates, with minimal manual floorplanning, over 2× reduction in normalized wirelength and more than 3× improvement in core density, while maintaining high cross-configuration robustness. Compared to the state-of-the-art VWR2A baseline, the proposed design exhibits superior scalability and routability. It thus provides a cost-effective, high-density, low-interconnect-overhead implementation pathway for Ångström-era DSIPs.
To address layout inefficiency, communication redundancy, unmeasurable power consumption, and limited scalability in Multi-Project Wafer (MPW) platforms for large-scale chip education and research, this paper proposes a high-density, low-cost, and scalable on-chip shared architecture. Methodologically: (1) an algorithm-driven automated floorplanning framework maximizes die area utilization; (2) a novel lightweight interconnect and resource-sharing mechanism leverages site-gap regions, eliminating redundant dedicated I/O and memory macros; (3) modular power-domain partitioning and on-die power monitoring enable per-project power characterization. Experimental results demonstrate up to 13× reduction in die area compared to conventional physically co-located MPW implementations, significantly improving resource utilization and project throughput—without requiring expertise in low-power ASIC design.
This paper addresses the automated layout design problem for stripboard (perfboard) circuits. We propose a declarative synthesis and multi-objective optimization approach based on Answer Set Programming (ASP). The problem is modeled holistically to enforce electrical connectivity, geometric constraints, and minimization of wire crossings. A two-stage solving strategy is adopted: first ensuring layout feasibility, then jointly optimizing board area and the number of component jumper connections (i.e., through-hole interconnections). Unlike traditional heuristic methods, our declarative formulation naturally encodes complex domain constraints, yielding higher-quality and more manufacturable solutions. Experimental evaluation across circuits of varying complexity demonstrates that our method consistently produces compact, low-crossing, and solder-friendly layouts—achieving an average 18.7% reduction in board area and a 32.4% decrease in jumper count. The approach is particularly suitable for electronic prototyping and educational applications.
To address the low efficiency of manual exploration and the difficulty of balancing performance gains against hardware overhead in RISC-V custom instruction design, this paper proposes CIDRE—a fully automated toolchain for end-to-end custom instruction synthesis, from dynamic hotspot analysis and pattern extraction to automatic generation of synthesizable nML hardware descriptions. Methodologically, CIDRE integrates a microarchitecture-aware instruction identification mechanism with an ASIP co-design flow, enabling joint evaluation of performance, area, and power. Evaluated on Embench and MiBench benchmarks, it achieves an average speedup of 1.83× (up to 2.47×) with custom instruction area overhead ≤24%, significantly outperforming existing manual or semi-automated approaches. Key contributions include: (1) the first end-to-end, microarchitecture-aware framework for automated custom instruction generation; (2) support for evaluatable and synthesizable nML output; and (3) empirical validation of the approach’s practicality and scalability under energy-efficiency and area constraints.
This study addresses the challenge of systematically comparing timing behavior of RISC-V processors across heterogeneous technology platforms—specifically, 20 nm FPGAs versus 7 nm FinFET ASICs. We propose a microarchitectural-level, cross-platform timing attribution methodology that integrates static timing analysis (STA), PVT-corner statistical characterization, and pipeline-stage decoupled modeling. Our approach establishes a three-component decomposition framework—logic, routing, and clock—and precisely localizes timing-critical transitions to individual pipeline stages. For the first time, we reveal that FPGA timing is dominated by routing parasitics and topology sensitivity, yielding wide yet scattered timing margins; in contrast, ASIC timing is governed by combinational logic depth and PVT stability, resulting in narrow, concentrated margins. Quantitatively, we identify the EX→MEM stage transition as the common critical path across both platforms. Based on this insight, we formulate predictive, heterogeneity-aware design guidelines for timing convergence.
This work addresses the limited reproducibility and comparability of machine learning research in electronic design automation (EDA), which stems from the absence of open, standardized datasets. To bridge this gap, the authors propose EDA-Schema-V2—the first standardized multimodal data schema encompassing the full EDA flow from logic synthesis to detailed routing. Leveraging open-source PDKs such as SkyWater 130nm and Nangate 45nm, along with the OpenROAD framework, they generate a large-scale open dataset comprising 7,776 design instances, over 275 million logic gates, and 36 million timing paths through systematic sweeps of process corners, clock periods, and placement parameters. The study defines twelve representative prediction tasks and establishes cross-stage predictability baselines, thereby providing a reproducible benchmark for ML-driven EDA research.
This work addresses the absence of standardized benchmarks for evaluating large language models (LLMs) and vision-language models (VLMs) in multi-stage optimization and collaboration with electronic design automation (EDA) tools within VLSI physical design. It introduces the first comprehensive evaluation framework tailored to this domain, comprising five dimensions—knowledge comprehension, report analysis, root-cause diagnosis, script generation, and end-to-end implementation—with 353 industry-validated questions verified by domain experts. The benchmark integrates real-world EDA environments, such as Cadence Innovus, enabling closed-loop assessment. Experimental results reveal that while current models perform reasonably on conceptual tasks, they exhibit significant deficiencies in tool interaction—evidenced by a mere 42.2% accuracy in Innovus script generation—and long-horizon reasoning. Incorporating human-in-the-loop workflows substantially enhances end-to-end design performance.
Early exploration of 3D integrated circuit architectures has been hindered by inaccurate modeling of layout-induced thermal, interconnect, and cache effects, leading to unreliable predictions of actual operating frequency and performance. This work proposes CLIP-3D, a left-shifted design flow that maps architectural configurations to physically aware representations without requiring signoff tools, and jointly optimizes cross-layer macro allocation and floorplanning through a thermal-aware placer. Its key innovation lies in introducing, for the first time, a closed-loop, analytical continuous-frequency model that directly targets actual BIPS (billion instructions per second) in joint optimization, eliminating reliance on manually weighted proxy objectives. By integrating McPAT, CACTI, and a HotSpot-compatible 3D thermal model, CLIP-3D substantially improves performance estimation accuracy and enables efficient architectural-level screening of designs free from thermal or interconnect-induced frequency throttling.
This work addresses the inefficiencies and semantic inconsistencies arising from separately implementing driver and monitor programs in traditional hardware module testing. To overcome this, the authors propose a domain-specific language (DSL) tailored to hardware communication protocols, which enables the unified specification of both driver and monitor logic through an imperative syntax, thereby ensuring their semantic consistency for the first time. Building upon this DSL, they develop a prototype tool that leverages waveform parsing and transaction-level trace inference techniques to accurately reconstruct protocol-compliant transaction sequences from raw signal waveforms. Experimental results demonstrate that the approach significantly improves development efficiency, with further validation planned on real-world interconnect protocols such as Wishbone and AXI-Stream.
At advanced technology nodes, the tight coupling between layout and electrical performance in analog circuits poses significant challenges for automated placement. This work proposes a row-height-quantized cell-based layout synthesis methodology, systematically introducing row-height quantization into analog circuit design for the first time. By optimizing row-height structures, modeling layout constraints, and enabling automatic mapping of analog modules onto quantized rows, the approach effectively bridges the performance gap between schematic and post-layout stages. Experimental results across multiple test cases demonstrate that the method achieves performance close to manual custom design, reducing the schematic-to-post-layout performance deviation by up to 68.5% and decreasing area overhead by as much as 24.1%.