BitIR: Cross-Architecture Fault Injection for Resilience Analysis of Heterogeneous GPU Applications

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear cross-backend discrepancies in fault behavior within heterogeneous GPU systems by proposing the first LLVM IR-level cross-architecture fault injection framework, which enables semantically equivalent injection of deterministic single-bit perturbations. Through large-scale benchmarking on the Polaris, Aurora, and Frontier supercomputers, this work reveals that different vendors exhibit markedly divergent responses to identical faults: Intel GPUs frequently mask errors yet are prone to hangs, NVIDIA GPUs tend toward hard failures, and AMD GPUs demonstrate workload-dependent behavior. These findings establish that resilience is not a backend-invariant property, underscoring the necessity of designing backend-aware and workload-aware fault tolerance strategies for heterogeneous high-performance computing environments.
📝 Abstract
Modern GPU-based HPC systems rely on heterogeneous vendor stacks, yet resilience studies are largely limited to single architectures, leaving it unclear how faults behave across different GPU backends. We present \emph{BitIR}, a cross-architecture fault injection framework that injects deterministic single-bit faults at the LLVM IR level, ensuring semantically equivalent perturbations prior to backend lowering and enabling direct cross-vendor comparison across NVIDIA, Intel, and AMD GPUs. We evaluate BitIR on three production supercomputers -- Polaris (ALCF, NVIDIA A100), Aurora (ALCF, Intel GPU Max 1550), and Frontier (OLCF, AMD Instinct MI250X) -- representing the full spectrum of current leadership-class GPU architectures. Using representative heterogeneous benchmarks, we conduct large-scale injection campaigns across all three systems and classify outcomes into masked results, silent data corruptions (SDCs), and failures. Our results show that identical faults produce markedly different behaviors across vendors: Intel most often masks faults but exhibits more hangs when faults escape masking, NVIDIA exposes more detectable hard failures, and AMD alternates between SDC-dominant and failure-dominant behavior depending on the benchmark and fault site. These findings demonstrate that resilience is not backend-invariant, underscoring the need for backend-aware and workload-aware fault mitigation.
Problem

Research questions and friction points this paper is trying to address.

cross-architecture fault injection
GPU resilience
heterogeneous computing
silent data corruption
vendor-specific behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Architecture Fault Injection
LLVM IR
Heterogeneous GPU
Resilience Analysis
Silent Data Corruption
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Maisy Dunlavy
University of Illinois Chicago, USA
M
Michael Papka
Argonne National Laboratory, USA
Zhiling Lan
Zhiling Lan
Professor of Computer Science, University of Illinois Chicago
cluster schedulingenergy efficiencyAI4Sysmodeling and simulationresilience