BLINK: Batch Normalization-based Integrity Checkpoints for In-Situ Detection and Mitigation of Diverse Weight Corruptions in DNN Accelerators

📅 2026-09-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
BLINK利用基于批归一化的在线检测和缓解方法,解决DNN加速器中由多种原因引起的权重损坏问题,无需主机干预或额外数据。
📝 Abstract
In safety-critical deployments, AI hardware must remain reliable against a broad spectrum of threats such as aging, soft errors, hard faults, and adversarial attacks (e.g. progressive bit flip attack (PBFA)). All of these corrupt stored weights while the chip keeps producing confident but inaccurate predictions. Detecting and mitigating such weight perturbations is crucial for safety-critical platforms. To that end, we propose BLINK, an on-chip batch normalization (BN)-based on-the-fly detection and mitigation approach, which is based on continual sensing of the shift in the activation statistics, and targets a wide variety of weight corruptions (random and localized faults as well as adversarial bit flips). BLINK operates in two phases: (1) off-line pre-characterization of the relationship of the activation shifts with inference accuracy drop, and (2) on-chip runtime detection and mitigation of weight corruptions. Upon detection, the flagged layer is re-centered to bring it closer to its stored clean reference within the same forward pass. BLINK is fully autonomous, eliminating the need for host communication, operation halts, or access to fine-tuning data. If the residual shift after mitigation indicates that accuracy has fallen below a user-set floor, a held-out watcher aborts the inference. Evaluated on ResNet-20/50 and MobileNetV2 for CIFAR-10/100, BLINK detects harmful corruptions with>99% precision across all fault types. Further, it recovers accuracy from 10% to 85.88% under 0.5% random bit flips (Resnet-50/CIFAR-10), up to 84% for localized faults (MobileNetV2/CIFAR-10), and from random-guess accuracy to 80%-83% under PBFA (ResNet-20/CIFAR-10). Hardware overhead estimates indicate that BLINK incurs negligible costs, with less than a 2% increase in latency and only a 0.53% increase in computation overhead.
Problem

Research questions and friction points this paper is trying to address.

AI hardware
weight corruptions
safety-critical platforms
adversarial attacks
aging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Batch Normalization
on-the-fly detection
weight corruptions
autonomous mitigation
activation statistics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marzia Khan
Purdue University
Akul Malhotra
Akul Malhotra
Research Scientist, IBM
In-memory computingAnalog AIDeep learning accelerators
S
Sumeet Kumar Gupta
Purdue University