🤖 AI Summary
Deploying ternary large language models (LLMs) on ternary compute-in-memory (TCiM) accelerators improves energy efficiency, but stuck-at-faults (SAFs) in memory cells cause significant accuracy degradation—especially under low-sparsity weight configurations. To address this, we propose Fault-Aware Symbol Transformation (FAST) coupled with bit-cell natural redundancy reprogramming: FAST dynamically corrects computation errors for ±1 weights, while redundancy reprogramming restores stuck-at-zero weights. Evaluated on BitNet b1.58 and Wikitext, our approach reduces perplexity by 35% under SAFs, incurring only <3% energy overhead, <7% latency penalty, and <1% area cost. This work introduces the first tightly integrated framework combining symbolic transformation with hardware-level redundancy reprogramming, establishing a novel, high-accuracy, low-overhead fault-tolerance paradigm for TCiM accelerators.
📝 Abstract
Ternary large language models (LLMs), which utilize ternary precision weights and 8-bit activations, have demonstrated competitive performance while significantly reducing the high computational and memory requirements of full-precision LLMs. The energy efficiency and performance of Ternary LLMs can be further improved by deploying them on ternary computing-in-memory (TCiM) accelerators, thereby alleviating the von-Neumann bottleneck. However, TCiM accelerators are prone to memory stuck-at faults (SAFs) leading to degradation in the model accuracy. This is particularly severe for LLMs due to their low weight sparsity. To boost the SAF tolerance of TCiM accelerators, we propose ReTern that is based on (i) fault-aware sign transformations (FAST) and (ii) TCiM bit-cell reprogramming exploiting their natural redundancy. The key idea is to utilize FAST to minimize computations errors due to SAFs in +1/-1 weights, while the natural bit-cell redundancy is exploited to target SAFs in 0 weights (zero-fix). Our experiments on BitNet b1.58 700M and 3B ternary LLMs show that our technique furnishes significant fault tolerance, notably 35% reduction in perplexity on the Wikitext dataset in the presence of faults. These benefits come at the cost of<3%,<7%, and<1% energy, latency and area overheads respectively.