Score
Designs and conducts analyses and experiments that measure, model, and quantify timing behavior and variations of microarchitectural components. This includes performing timing analysis to evaluate timing-based effects and exploitability, model read-disturbance or other timing phenomena, and compare performance-related tradeoffs such as area and frequency.
To address the low efficiency and difficulty in root-cause localization during pre-signoff MCMM timing debugging in VLSI design, this paper proposes a multi-LLM collaborative intelligent agent system. Methodologically, it introduces (1) the Timing Debugging Relation Graph (TDRG)—the first domain-specific knowledge graph integrating circuit topology and timing constraints; (2) an Agentic RAG framework unifying graph-based retrieval, executable code reasoning, and hierarchical planning; and (3) an end-to-end pipeline for automated report parsing, root-cause identification, and repair recommendation generation. Evaluated on industrial-scale benchmarks, the system achieves 98% success rate on single-report debugging and 90% on multi-report joint debugging, substantially reducing debug turnaround time. This work represents the first systematic adoption of embodied intelligent agents in VLSI timing verification, establishing a new paradigm for AI-driven signoff automation.
This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.
This study addresses the challenge of systematically comparing timing behavior of RISC-V processors across heterogeneous technology platforms—specifically, 20 nm FPGAs versus 7 nm FinFET ASICs. We propose a microarchitectural-level, cross-platform timing attribution methodology that integrates static timing analysis (STA), PVT-corner statistical characterization, and pipeline-stage decoupled modeling. Our approach establishes a three-component decomposition framework—logic, routing, and clock—and precisely localizes timing-critical transitions to individual pipeline stages. For the first time, we reveal that FPGA timing is dominated by routing parasitics and topology sensitivity, yielding wide yet scattered timing margins; in contrast, ASIC timing is governed by combinational logic depth and PVT stability, resulting in narrow, concentrated margins. Quantitatively, we identify the EX→MEM stage transition as the common critical path across both platforms. Based on this insight, we formulate predictive, heterogeneity-aware design guidelines for timing convergence.
Existing static performance analysis lacks lightweight, platform-agnostic metrics for compile-time performance estimation. Method: We propose memory access count (MEMS) as a static performance proxy, implemented via Clang AST rewriting and source-level automatic instrumentation to embed path-wise MEMS modeling and counting logic directly into the code. Contribution/Results: We systematically validate MEMS across ten classical algorithms. Experiments reveal— for the first time—that MEMS exhibits strong positive correlation with runtime performance *within* a given program across different execution paths, but weak correlation *across* distinct programs; this delineates MEMS’s practical applicability boundary. Crucially, MEMS incurs zero runtime overhead and requires no hardware-specific support, offering a portable, low-cost metric for static performance prediction. By bridging the gap between compile-time analysis and empirical performance behavior, MEMS significantly enhances the practicality of compiler optimizations and static performance modeling.
Microarchitectural timing channels enable implicit cross-security-boundary information leakage, undermining temporal isolation guarantees in secure systems. Method: This paper proposes the first full temporal isolation scheme for RISC-V, centered on an ISA-native timing fence instruction `fence.t` and a hardware-level full-state zeroing mechanism that systematically clears non-architectural core state to ensure history-independent context-switch latency bounds. Contribution/Results: We formalize the RISC-V ISA extension, adapt the seL4 microkernel, and implement the scheme in the open-source CVA6 processor. The solution eliminates all major on-core timing channels—including cache, branch predictor, and TLB-based channels—while incurring less than 1% performance overhead and negligible hardware cost. Crucially, it provides formally verifiable temporal isolation guarantees, establishing a foundation for high-assurance real-time and security-critical systems.
本文针对多核系统中的硬件干扰问题,提出了一种基于好奇心驱动探索算法的方法,以更高效地识别和分析干扰源。
This work addresses a long-overlooked yet critical vulnerability in MIPS embedded processors: severe cross-core microarchitectural side-channel leakage when simultaneous multithreading (SMT) is enabled. The authors propose MIPSBLEED, a novel framework that systematically uncovers timing-based information leakage across cores through the L1 data and instruction caches as well as the execution engine. By combining assembly-level probing, microarchitectural timing modeling, and quantitative leakage analysis, they construct a high-resolution, single-trace attack that requires no privileged access. Demonstrated on real hardware, the attack successfully recovers elliptic curve cryptography keys, confirming substantial information leakage from security-critical components. These findings underscore the urgent need for lightweight yet effective isolation mechanisms in SMT-enabled MIPS architectures.
本文通过分析SPEC CPU 2026在AMD EPYC 9755处理器上的性能特征,采用多角度方法识别系统瓶颈和工作负载行为差异。
研究通过结合概率设计时延分析与Kuksa实现监控,解决了软件定义车辆中因中间件通信导致的时间不确定性问题。
Existing architectural simulators struggle to uncover the complex causal relationships between microarchitectural events and program behavior, making it difficult to attribute performance bottlenecks across abstraction layers. This work proposes Microflow, an observability framework that treats causality as a first-class analytical construct. Microflow introduces MFIR, an intermediate representation that explicitly encodes software-hardware dependencies, enabling queryable causal inference, revelation of latent phenomena, and precise critical path decomposition. By integrating counterfactual analysis with cross-layer dependency modeling, Microflow successfully identifies hidden bottlenecks on SPEC CPU 2017 benchmarks that are missed by conventional approaches—such as implicit misprediction overhead in leela and inter-loop resource contention in mcf.