Score
Designs, implements, and analyzes the organization and microarchitectural components of computing systems—processors, pipelines, caches and memory hierarchies, interconnects, caches/coherence, instruction-set and accelerator interfaces—along with the hardware–software interfaces that control them. Evaluates and optimizes trade-offs among performance, power, area, cost, and reliability using models, simulations, prototypes, and hardware implementations.
Facing slowing transistor scaling, rising power consumption, and intensifying instruction-level parallelism (ILP) bottlenecks, this work systematically analyzes the evolution of multicore CPU architectures across four dimensions: microarchitectural design, cache coherence, OS scheduling co-design, and quantitative evaluation. We propose the first full-stack analytical framework integrating hardware microarchitectural modeling, system software adaptation, and multi-tiered benchmarking (e.g., SPEC CPU, PARSEC). The study clarifies the dual drivers—“power wall” and “ILP wall”—behind the industry’s shift to multicore processors and introduces a scalable, joint performance–energy-efficiency evaluation methodology. Our framework enables rigorous, cross-layer analysis of architectural trade-offs and provides theoretical foundations and methodological guidance for heterogeneous integration, energy-efficient design, and hardware–software co-optimization.
To address critical challenges in SoC design—including ambiguous system-level modeling semantics, poor interoperability across heterogeneous computational models (e.g., dataflow and neural networks), and the decoupling of design-space exploration from verification—this paper proposes a co-communication mechanism ensuring semantic consistency across multiple models. The approach establishes an integrated toolchain supporting system-level modeling, simulation-driven verification, hardware-software co-design space exploration, and joint power-performance analysis. Innovatively, it unifies dataflow modeling with system-level abstractions to enable functional correctness verification and quantitative energy-efficiency evaluation for representative applications such as video processing and AI acceleration. Experimental results demonstrate that the methodology significantly improves early-stage SoC design iteration efficiency and enhances the reliability of architectural decision-making.
Modern heterogeneous architectures—including multi-core CPUs, TPUs, RipTide, and Catapult—face a fundamental trade-off among energy efficiency, latency, and hardware flexibility, especially amid evolving domain-specific design trends. Method: This paper proposes a three-dimensional trade-off-driven reconfigurable computing architecture optimization framework. Leveraging systematic cross-layer modeling and performance evaluation, it unifies major accelerator design paradigms for the first time and introduces a heterogeneous computational model supporting dynamic hardware reconfiguration. Contribution/Results: Experimental evaluation demonstrates that the framework improves computational efficiency by 30–50%, reduces dynamic power consumption by over 40%, and effectively breaks the traditional fixed-architecture bottleneck in the energy-efficiency–latency–flexibility triad. The work delivers a theoretically grounded, implementation-ready framework and concrete architectural design guidelines for next-generation adaptive heterogeneous computing systems.
Conventional architecture design relies heavily on manual expertise, suffers from siloed hardware-software optimization, and incurs prohibitively high costs for design-space exploration. Method: This paper proposes LPCM, an LLM-driven three-tier collaborative framework that establishes the first “human-agent-model” co-design paradigm, deeply embedding large language models into a closed-loop hardware-software co-design workflow to overcome limitations of single-stage and fragmented optimization. It innovatively integrates 3D Gaussian splatting–based workload modeling with system-level co-design methodology. Contribution/Results: At Level 1 validation, LPCM achieves full automation of the end-to-end architecture design pipeline. Experiments demonstrate substantial reduction in design cycle time and human effort, establishing a scalable, reusable technical pathway toward fully autonomous, full-stack chip design.
This study investigates the performance and energy efficiency differences between ARM and x86-64 laptop processors, demonstrating that these disparities stem not only from instruction set architecture (ISA) but also significantly from system-level design choices. For the first time, the authors conduct a comprehensive evaluation on real-world laptop platforms—Apple M3 and AMD Ryzen 7 3750H—combining fine-grained power measurements with microarchitectural analysis. Using assembly-level benchmarks (recursive Fibonacci, integer matrix multiplication), cross-platform performance counters, and portable C-based probes, they systematically assess the impact of architectural and integration factors on energy efficiency. Results reveal that while the Ryzen platform excels in branch-intensive workloads, the Apple platform achieves substantially superior energy efficiency, reducing energy-per-computation by 5.82× and 6.38× respectively, thereby highlighting the critical role of non-ISA design elements.
Existing methodologies—such as SPEC CPU2017, Design of Experiments (DoE), and Randomized Controlled Trials (RCTs)—struggle to accurately attribute overall system performance to individual hardware components due to their inability to effectively isolate component-level contributions, resulting in substantial evaluation variability (SPEC score deviations ranging from 12.16% to 436.80%). This work proposes a novel methodology that integrates controlled experimentation with a theoretical attribution model, enabling, for the first time, precise and stable attribution of system performance to specific hardware components. The proposed approach significantly outperforms conventional techniques, offering high cost-effectiveness while overcoming inherent limitations in component evaluation and system design present in current practices.
Traditional computer architecture simulation suffers from poor scalability, limited reproducibility, and excessive customization due to its reliance on implicit scripts and directory conventions. This work proposes the first end-to-end explicit and service-oriented simulation framework, which models hardware topologies declaratively via graph representations, automatically generates executable simulation code, and employs a stateless runner for automated task scheduling and structured result management. The approach eliminates the need for manual simulation programming and enables systematic exploration through automatic expansion of configuration–benchmark matrices. Evaluated across 96 GPU workloads, the framework achieves a median kernel time error of only 0.18% compared to hand-tuned MGPUSim configurations—covering 95.8% of all configurations—with a negligible per-simulation overhead of just 1.6 seconds.
This work addresses the lack of systematic educational resources in high-performance computing (HPC) networking, which poses a significant barrier for researchers entering the field. It presents the first comprehensive integration of the HPC networking stack, covering communication protocol layers, programming interfaces such as MPI, control plane mechanisms, high-speed interconnect technologies, and custom link-layer hardware. The exposition is anchored by a detailed case study of the El Capitan supercomputer architecture at Lawrence Livermore National Laboratory. By offering a well-structured, practice-oriented primer, this contribution fills a critical educational gap and substantially lowers the entry barrier for researchers seeking to master core HPC networking technologies.
This study systematically examines the four-decade evolution of synchronization mechanisms and in-network computing architectures in large-scale parallel systems. Tracing the trajectory from the NYU Ultracomputer to modern exascale supercomputers, it integrates key technological milestones—including Fetch-and-Add, multistage interconnection networks, MPI, PCIe atomics, GPU cache coherence mappings, and HIP/Triton compilation stacks—to uncover, for the first time, the dynamic interplay and competition among shared-memory, message-passing, and in-network computing paradigms. The work elucidates the continuous co-evolution of synchronization primitives across hardware-software boundaries, offering critical historical context and architectural insights for the design of future heterogeneous supercomputing systems.