🤖 AI Summary
This study addresses the unclear cross-backend discrepancies in fault behavior within heterogeneous GPU systems by proposing the first LLVM IR-level cross-architecture fault injection framework, which enables semantically equivalent injection of deterministic single-bit perturbations. Through large-scale benchmarking on the Polaris, Aurora, and Frontier supercomputers, this work reveals that different vendors exhibit markedly divergent responses to identical faults: Intel GPUs frequently mask errors yet are prone to hangs, NVIDIA GPUs tend toward hard failures, and AMD GPUs demonstrate workload-dependent behavior. These findings establish that resilience is not a backend-invariant property, underscoring the necessity of designing backend-aware and workload-aware fault tolerance strategies for heterogeneous high-performance computing environments.
📝 Abstract
Modern GPU-based HPC systems rely on heterogeneous vendor stacks, yet resilience studies are largely limited to single architectures, leaving it unclear how faults behave across different GPU backends. We present \emph{BitIR}, a cross-architecture fault injection framework that injects deterministic single-bit faults at the LLVM IR level, ensuring semantically equivalent perturbations prior to backend lowering and enabling direct cross-vendor comparison across NVIDIA, Intel, and AMD GPUs. We evaluate BitIR on three production supercomputers -- Polaris (ALCF, NVIDIA A100), Aurora (ALCF, Intel GPU Max 1550), and Frontier (OLCF, AMD Instinct MI250X) -- representing the full spectrum of current leadership-class GPU architectures. Using representative heterogeneous benchmarks, we conduct large-scale injection campaigns across all three systems and classify outcomes into masked results, silent data corruptions (SDCs), and failures. Our results show that identical faults produce markedly different behaviors across vendors: Intel most often masks faults but exhibits more hangs when faults escape masking, NVIDIA exposes more detectable hard failures, and AMD alternates between SDC-dominant and failure-dominant behavior depending on the benchmark and fault site. These findings demonstrate that resilience is not backend-invariant, underscoring the need for backend-aware and workload-aware fault mitigation.