🤖 AI Summary
This study addresses the execution distortion and high overhead introduced by existing GPU kernel tracing tools, whose pre-compilation instrumentation disrupts compiler optimizations. We propose Xtrace, a system that pioneers zero-compilation-interference binary-level instruction stitching to achieve high-fidelity intra-kernel GPU tracing. By repurposing dead-value registers, parsing compiler hazard tables, and optimizing instruction scheduling alongside control bits, Xtrace supports multiple architectures while minimizing runtime overhead. Experimental results demonstrate that Xtrace preserves 94–98% of original instructions with a marginal tracing overhead of only 0.9–2.8%. Furthermore, it successfully accelerates the optimization iteration cycle for LLM kernels, yielding a 5.2–13.3% throughput improvement in FlashAttention.
📝 Abstract
Modern GPU kernels fuse increasingly more work into a single kernel, and intra-kernel tracing has become the mainstream method to profile them. Tracing inserts probes into the kernel to record its runtime states, and the fidelity of the trace determines the efficiency of performance optimization. Unfortunately, existing tools insert probes before compilation. These tools interfere with the compiler's optimizations, so they trace a different binary from the one the GPU executes. They also add significant runtime overhead.
Xtrace is the first GPU kernel tracing system with near-zero compile-time interference and minimized runtime overhead. Xtrace inserts probes directly into the compiled kernel binary. It reuses only the registers that hold dead values at the insertion address and resolves all hazards with the compiler's hazard tables. It further schedules the instruction order, register allocation, and control bits to minimize the runtime overhead the probe introduces. Xtrace supports 19 NVIDIA and AMD GPU architectures, and is publicly available for use at https://g-watch.github.io.
We evaluate Xtrace on major production large language model (LLM) kernels against the state-of-the-art tracers Neutrino and IKET from NVIDIA. On H100, B300, and MI300X GPUs, Xtrace preserves 94-98% of the instructions of the kernel, while existing tools preserve only 8-48%. Xtrace adds only 0.9-2.8% overhead, while existing tools add 3.8-75.6%. Xtrace guides a coding agent to reach the same FlashAttention-3 performance with 3.9x fewer iterations than existing traces do. Thanks to our binary-level instrumentation, Xtrace also traces the faster closed-source cuDNN kernel, which guides the agent to lift the open-source FlashAttention-4 by 5.2-13.3% in throughput.