Score
Designs, implements, and applies tools and methods to measure and analyze memory and data-transfer throughput during device or runtime execution and to identify bandwidth-related bottlenecks; and develops concrete optimizations—e.g., data layout and tiling, prefetching, scheduling, concurrency control, DMA/bulk-transfer tuning, compression, and runtime throttling—to increase effective memory bandwidth and reduce stalls.
To address the challenge of optimizing data placement in HBM+DDR heterogeneous memory systems, this paper proposes a lightweight, non-intrusive, application-level memory analysis and control framework. Leveraging a detailed memory subsystem model on Intel Sapphire Rapids platforms—integrated with empirical bandwidth/latency measurements, runtime allocation tracing, and policy injection—the work quantifies, for the first time, the performance inflection point of HBM-DRAM co-utilization: retaining only 60–75% of critical data in HBM achieves 90% of the platform’s peak performance. The approach requires no source-code modification or recompilation. Evaluated across multiple benchmarks, it significantly improves performance for memory-intensive applications while reducing HBM resource consumption by over 30%. This establishes a deployable, fine-grained data placement optimization paradigm for heterogeneous memory systems.
Hardware cache prefetching, memory scheduling, and channel interleaving obscure program context, hindering context-aware memory management. Method: This paper proposes a lightweight context-aware memory system that encodes program execution markers and object address ranges—i.e., contextual state—directly into standard memory read address streams as detectable metadata packets, requiring no privileged access or custom drivers. It integrates metadata injection, HMU-based telemetry hardware, and near-memory computing to enable bidirectional embedding and parsing of context within the address stream. Contribution/Results: A prototype demonstrates highly reliable metadata decoding from real address traces, enabling runtime fine-grained data scheduling, priority-aware memory management, and dynamic reconfiguration of memory devices. To our knowledge, this is the first end-to-end verifiable, zero-intrusion, fully user-space context-aware memory architecture.
This study addresses the absence of open standards for CPU pipeline visualization tools and the difficulty in localizing performance bottlenecks. To this end, it proposes an open-source event stream format alongside Catscan, an interactive viewer. Methodologically, this work introduces a structured event stream based on transactional relationships, integrating typed event modeling, persistent highlighting techniques, and domain-specific search algorithms to enable microarchitectural trace analysis from symptoms down to individual instructions. Furthermore, it supports resource-oriented views synchronized with comparative trace alignment. By successfully reproducing industry-grade debugging workflows, this project provides the community with production-validated microarchitectural visualization infrastructure.
Modern processors are often limited by memory subsystem performance rather than computational throughput, yet existing benchmarks inadequately capture the interplay between memory accesses and compute instructions in determining real-world throughput. To address this, we design and implement Arm-membench—the first fine-grained memory bandwidth benchmark for the Armv8 architecture, fully supporting all current SIMD extensions—and achieve the first complete port and deep optimization of x86-membench to Arm. Leveraging Armv8 ISA modeling, custom SIMD assembly micro-benchmarks, and multi-core microarchitectural analysis, we uncover a previously unreported bottleneck: instruction fetch and decode width becomes critical under memory-intensive workloads. Empirical evaluation on Fujitsu A64FX, Ampere Altra, and Cavium ThunderX2 demonstrates that Arm-membench precisely quantifies pipeline-level bottlenecks under cache-bandwidth constraints, providing a novel methodology for performance modeling and microarchitectural analysis on Arm platforms.
DRAM bandwidth bottlenecks caused by indirect memory accesses—exacerbated by limited DRAM controller visibility, small request buffers, and insufficient memory-level parallelism—hinder modern multi-core systems. To address this, we propose DX100, a shared, programmable data-access accelerator. DX100 introduces the first general-purpose ISA-compatible architecture for hardware-accelerated indirect memory access, offloading batched indirect address computation and memory requests from multiple cores. It enables cross-core sharing, dynamic request reordering, interleaving, and merging—significantly improving row-buffer hit rates and DRAM bandwidth utilization. A fully automated MLIR-based compilation flow enables zero-modification porting of existing applications. Evaluated on 12 benchmarks spanning scientific computing, databases, and graph analytics, DX100 achieves a 2.6× speedup over a multi-core baseline and outperforms the state-of-the-art indirect prefetcher by 2.0×.
This study addresses the "instruction wall" bottleneck wherein increased GPU memory bandwidth fails to yield proportional acceleration in database queries. We introduce Valk, a multi-source performance analysis tool integrating multiple profilers with TPC-H benchmarks for cross-hardware diagnosis. Our investigation reveals the compute-bound mechanisms underlying high-bandwidth scenarios, elucidating why the GH200 achieves only a 5.2× speedup over the L4. Accordingly, we propose three optimization strategies: cache optimization, enhanced parallelism, and instruction streamlining. Experimental results demonstrate that these approaches effectively overcome the instruction wall limitation, fully unleashing the data processing potential of next-generation GPUs. Ultimately, this work provides both theoretical foundations and practical guidance for performance tuning in heterogeneous computing environments.
This work addresses the problem of excessive and ineffective hardware prefetching in datacenter workloads, which wastes precious memory bandwidth. The authors propose a novel hardware-software cooperative prefetching mechanism that leverages page table entries to convey page-level prefetch hints, enabling dynamic control over hardware prefetcher behavior without requiring modifications to the instruction set architecture or application binaries. The approach is compatible with existing state-of-the-art prefetchers—such as BOP, SPP+PPF, and Pythia—and selectively disables prefetching on non-critical pages at runtime to balance coverage and bandwidth efficiency. Experimental evaluation demonstrates that the proposed technique reduces ineffective prefetch requests by approximately 40% under representative datacenter workloads and achieves performance improvements of up to 4.1%.
This work addresses the limitations of traditional hardware prefetchers, which rely on manual analysis of program traces and struggle to handle performance anomalies in real-world workloads. The authors propose a performance-anomaly-driven automated design methodology that employs an intelligent agent to diagnose prefetching failures, automatically synthesize a Mixture-of-Prefetchers (MoP), and continuously refine it through simulation. This approach represents the first practical, agent-driven framework capable of generating synthesizable RTL for prefetchers, outperforming state-of-the-art handcrafted designs on previously unseen workloads. Evaluated on SPEC CPU2006/2017, MoP achieves a 61.1% geometric mean IPC improvement over no prefetching and surpasses Alecto, Berti, and Pythia by 14.5%–23.6%. Its RTL implementation in a 6nm process occupies only 110KB of storage and 0.0347mm² of silicon area.
This work addresses the inefficiency of traditional approaches to predicting workload performance under varying memory configurations, which typically rely on time-consuming simulations or repeated measurements. The study reveals, for the first time, a predictable relationship between cycles per instruction (CPI) and maximum memory stall across diverse workloads. By leveraging hardware performance counters collected from a single native execution—combined with mechanistic insights and empirical data—the authors construct a regression model that enables highly accurate, simulation-free first-order performance prediction. Evaluated across six machine configurations and two simulators, the method reduces CPI prediction error by 2× compared to the best existing single-run techniques. On ARM servers, it achieves a median error of 12.7% and a 90th-percentile error of 35.9%, maintaining robust accuracy even when extrapolating to memory latencies up to 8× higher than baseline.
研究通过结合静态代码特征、动态执行轨迹和内核级资源数据,发现13种性能原型,并提出一个多信号回归检测框架,以提高软件性能分析和预测的准确性。