🤖 AI Summary
This study addresses the degradation of DNN inference accuracy caused by hardware faults in heterogeneous edge accelerators, a critical issue overlooked by existing partitioning methods that neglect runtime reliability. To bridge this gap, this work pioneers the integration of runtime reliability into multi-objective partitioning optimization. By analyzing layer-wise sensitivity through system-level fault injection and leveraging the NSGA-II algorithm alongside Eyeriss/SIMBA simulation platforms, we propose a three-tier-aware partitioning framework that jointly optimizes latency, energy consumption, and robustness. Experimental results demonstrate that, with only 2–3% additional overhead, the proposed framework improves Top-1 accuracy by 9.81% and reduces fault-induced accuracy loss by 36.86%. These findings effectively resolve the longstanding challenge of fault-unaware partitioning in traditional approaches, offering a practical solution for reliable DNN deployment on heterogeneous edge systems.
📝 Abstract
Deep neural networks (DNNs) deployed on heterogeneous edge accelerators are increasingly exposed to runtime hardware faults. Tight power and resource budgets push these platforms toward aggressive voltage scaling and limit the use of hardware fault mitigation, while process variation, thermal stress, and device aging raise fault rates further. The resulting transient bit level corruptions propagate through inference and can sharply reduce prediction accuracy. Existing DNN partitioning frameworks optimize latency and energy under fault free assumptions, so their deployment strategies often fail to sustain reliable inference in realistic operating conditions. This paper presents RAPID DNN, a reliability aware partitioning framework that adds runtime reliability as a third optimization objective alongside latency and energy. RAPID DNN first characterizes layer wise reliability through systematic fault injection in the weight and activation domains to identify vulnerability critical layers. The resulting sensitivity profile guides a three objective NSGA II optimizer that jointly minimizes inference latency, energy consumption, and expected fault induced accuracy degradation. We evaluate RAPID DNN on AlexNet, SqueezeNet, ResNet18, VGG16, and MobileNetV2 using heterogeneous Eyeriss and SIMBA accelerator profiles. Across fault models and injection probabilities, it consistently improves inference robustness over the fault agnostic baseline. Under the representative mixed random fault configuration, RAPID DNN raises average Top 1 accuracy by 9.81 percentage points and reduces average expected fault induced accuracy degradation by 36.86%, at a cost of 2.38% average latency overhead and 3.34% average energy overhead. These results show that treating reliability as a partitioning objective enables robust DNN deployment on heterogeneous edge accelerators under runtime hardware faults.