Hardware-Attributed Operator Profiling for PyTorch

📅 2026-07-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of framework profilers lacking hardware counters, GPU profilers missing operator attribution, and the error-prone nature of manual correlation by proposing an automated hardware attribution pipeline. Methodologically, it integrates CUPTI, NVTX interval trees, and Inductor debug artifacts to design three complementary attribution paths alongside a call-order matching algorithm, incorporating deduplication and clock-locking mechanisms to ensure data consistency. Experimental results demonstrate that this approach achieves 95%–100% kernel runtime attribution for compilation workloads on Blackwell GPUs, enabling FX graph optimizations that yield up to a 2.24× speedup. This work provides a scalable solution for fine-grained performance bottleneck analysis.
📝 Abstract
Framework profilers expose operator timing without hardware counters; GPU profilers expose hardware counters without operator attribution. Bridging this gap manually is error-prone and does not scale. We present Operator Profiler, a hardware attribution pipeline that automatically links hardware metrics to PyTorch operators via three complementary attribution paths: torch.profiler CUPTI correlation, NVTX temporal enclosure with per-stream interval trees, and Inductor fusion- map enrichment from debug artifacts. NVIDIA Nsight Compute (ncu) hardware counters are matched to NVIDIA Nsight Systems (nsys) kernel records via invocation-order matching, avoiding timestamp joins across incompatible clock domains. A curated 20-counter metric set with duration-weighted aggregation covers all hardware bottleneck axes, layer deduplication reduces ncu replay time by a factor of N/K for models with N layers across K unique structural classes, and GPU clock locking controls the kernel-duration aggregates used for operator-level comparison. On an NVIDIA RTX PRO 6000 Blackwell, Operator Profiler attributes 95-100% of kernel runtime for compiled workloads (GPT-2, SDPA Attention); black-box library backends such as cuDNN RNN are correctly surfaced as greater than 85% unattributed rather than silently dropped. Applied to profile-guided FX graph optimization, attributed profiles yield 1.76x-2.24x profiled-kernel-time speedups on the two compiled optimization case studies; a third LSTM diagnostic case identifies cuDNN re-dispatch as a structural fix rather than an FX graph rewrite.
Problem

Research questions and friction points this paper is trying to address.

Operator Profiling
Hardware Counters
PyTorch
Performance Attribution
GPU Profiling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hardware-Attributed Operator Profiling
CUPTI Correlation
NVTX Temporal Enclosure
Profile-Guided Optimization
Inductor Fusion-Map
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Logan Chu
Yotta Labs
D
Dong Li
Yotta Labs