Score
Designs and validates analytical and empirical models that predict and bound latency and throughput for software and hardware systems, explicitly capturing microarchitectural effects such as cache behavior and operator fusion. Builds and tunes latency-aware components and optimizations — including schedulers, routers, priority queues, low-latency networking stacks, and mappings of algorithms to hardware engines — and produces multi-fidelity performance analyses to inform latency-aware routing, scheduling, and implementation decisions.
This work addresses the challenges posed by large language models to many-core architectures, particularly the demand for high parallelism and on-chip memory capacity, which render traditional RTL simulation inefficient for accurately modeling latency-sensitive interconnect behavior. The authors propose an end-to-end rapid modeling methodology that combines hierarchical interconnect modeling, timing-abstracted scratchpad memory (SPM) access, and NoC router remapping to preserve critical timing details while significantly simplifying non-essential hardware components. Evaluated on the TeraNoC platform—a 1024-core system with 4 MiB shared L1 SPM—the approach achieves up to 115× simulation speedup with less than 7% error, enables fine-grained performance profiling, and successfully guides optimizations in FlashAttention-2 to mitigate interconnect stalls. Furthermore, design space exploration of the NoC yields substantial throughput improvements.
To address the prohibitively slow speed of cycle-accurate simulators (e.g., gem5) in microarchitectural design space exploration, this paper proposes Concorde—a CPU performance modeling framework that synergistically integrates component-wise analytical modeling with lightweight machine learning. Its key innovation lies in the first use of interpretable, analytically derived performance distributions—characterizing caches, pipelines, and branch predictors—as input features to drive distribution-aware representation learning and efficient regression for program-level CPI prediction. Compared to gem5, Concorde achieves >10⁵× speedup with only ~2% mean absolute CPI error. It enables 150 million design evaluations within one hour and supports fine-grained, cross-program and cross-microarchitecture performance attribution. By unifying analytical insight with data-driven generalization, Concorde overcomes the longstanding accuracy-efficiency trade-off inherent in both traditional simulation and purely empirical modeling approaches.
This work addresses the performance degradation of general-purpose cores in heterogeneous SoCs caused by stringent deadline requirements of hardware accelerators sharing the last-level cache. To mitigate this issue, the authors propose HyDRA, a novel dynamic cache management mechanism that, for the first time, models the unique cache reuse behavior of accelerators. HyDRA integrates clustering-based LERN reuse prediction, deadline-aware scheduling, and dynamic bypass decisions to simultaneously meet accelerator deadlines and optimize system throughput. Experimental results demonstrate that HyDRA significantly improves overall performance across diverse workloads and accelerator configurations while substantially reducing deadline miss rates.
This work addresses the lack of systematic evaluation of existing hardware priority queue architectures on modern platforms, which has hindered the balanced optimization of performance, resource overhead, and scalability. For the first time, it comprehensively reimplements and benchmarks multiple classic hardware priority queue designs on contemporary FPGA platforms, establishing an open-source framework for testing and analysis. The study provides a quantitative assessment of these architectures in terms of latency, resource utilization, and scalability. By filling a long-standing gap in systematic benchmarking, this research delivers empirical insights and an open foundation to guide the development of future high-performance priority queue implementations.
This work challenges the conventional focus on core utilization in resource management, which often overlooks the practical performance constraints imposed by power and thermal limits in modern multicore processors. Instead, it proposes a new paradigm centered on power budgeting, elevating idle-core waiting strategies to first-class design considerations. Rather than aggressively reclaiming idle cores—a practice that frequently overestimates benefits and incurs substantial scheduling overhead—the approach leverages efficient waiting mechanisms to release redistributable compute capacity. Empirical analysis on AMD EPYC platforms, accounting for processor topology, idle duration, and waiting policies, demonstrates that such strategies achieve a superior trade-off between energy efficiency and performance, offering greater practical advantages in real-world systems.
This study systematically evaluates the reliability of four machine learning–based ranking models—NeuroScalar, SimNet, Concorde, and OneDSE—in ordering hardware configurations at the program phase level for microarchitectural design space exploration. Across structural parameter and behavioral policy scenarios, the analysis—integrating cycle-accurate simulation, Bayesian accuracy assessment, and information-theoretic methods—reveals, for the first time, that a substantial fraction (22.4%) of program windows exhibit counterintuitive rankings in structural settings, and inter-model consistency remains low (23.3%–39.9%). In behavioral policy scenarios, most models fail to surpass a featureless baseline, with the best achieving only a 2.1-percentage-point improvement. The work further establishes a theoretical upper bound on ranking accuracy when critical microarchitectural states are unobservable, demonstrating inherent limitations of instruction-stream–based approaches.