CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity Windows

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing helper-thread prefetching approaches struggle to balance efficiency and overhead across diverse workloads and processors due to their reliance on fixed or exhaustive granularity selection. This work proposes an automatic prefetch granularity selection mechanism constrained by cache capacity: it estimates the number of cache lines filled per iteration in target code regions through offline analysis, constructs a capacity-aware window to determine the optimal prefetch block size, and enforces block-level synchronization to bound the helper thread’s execution lookahead. By integrating cache capacity budgeting directly into granularity decisions—eliminating online search—the method achieves near-optimal performance with low overhead. Evaluated on Intel and AMD platforms, it delivers average speedups of 1.54× and 1.33×, respectively, outperforming Ghost Threading by 15.8% and 10.8%, while capturing over 99% of the optimal performance within the candidate set.
📝 Abstract
Helper-thread prefetching hides the latency of irregular memory accesses by executing address dependency chains ahead of the main thread. However, its effectiveness depends on the range of future iterations covered by the helper thread. A fixed coverage range cannot consistently accommodate different workloads and processors, whereas exhaustively evaluating candidate configurations incurs substantial configuration cost. This paper presents CW-Ghost, which uses a single offline profiling run to estimate the average demand cache line fill volume generated per target iteration in a target region. CW-Ghost combines this estimate with a cache capacity budget to derive a Capacity Window, which determines the iteration granularity of each prefetch chunk. In addition, bounded chunk-level synchronization limits the number of chunks by which the helper thread may run ahead of the main thread. Across 14 workload instances evaluated on Intel and AMD CPU platforms, CW-Ghost achieves geometric mean speedups of 1.54x and 1.33x, respectively, over the original programs. Compared with Ghost Threading, it improves geometric mean performance by 15.8% and 10.8%, respectively, while achieving more than 99% of the empirically optimal performance within the candidate set on both platforms. These results demonstrate that cache capacity constraints can effectively guide the selection of granularity for helper-thread prefetching.
Problem

Research questions and friction points this paper is trying to address.

helper-thread prefetching
iteration granularity
cache capacity
memory latency
configuration cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

helper-thread prefetching
capacity window
granularity selection
cache-aware prefetching
search-free optimization
🔎 Similar Papers
2024-10-04arXiv.orgCitations: 1
Ya Zhang
Ya Zhang
Shanghai Jiao Tong University
Machine learningComputer visionMedical Imaging
T
Tong Lei
Laboratory of Digitizing Software for Frontier Equipment, National University of Defense Technology, Changsha 410073, China; National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Y
Yao Chen
Laboratory of Digitizing Software for Frontier Equipment, National University of Defense Technology, Changsha 410073, China; National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Y
Yonggang Che
Laboratory of Digitizing Software for Frontier Equipment, National University of Defense Technology, Changsha 410073, China; National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
C
Chuanfu Xu
Laboratory of Digitizing Software for Frontier Equipment, National University of Defense Technology, Changsha 410073, China; National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
H
Haozhong Qiu
Laboratory of Digitizing Software for Frontier Equipment, National University of Defense Technology, Changsha 410073, China; National Key Laboratory of Parallel and Distributed Computing, College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Yusong Tan
Yusong Tan
National University of Defense Technology
computeroperating systemcloudAI