ReTern: Exploiting Natural Redundancy and Sign Transformations for Enhanced Fault Tolerance in Compute-in-Memory based Ternary LLMs

📅 2025-06-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Deploying ternary large language models (LLMs) on ternary compute-in-memory (TCiM) accelerators improves energy efficiency, but stuck-at-faults (SAFs) in memory cells cause significant accuracy degradation—especially under low-sparsity weight configurations. To address this, we propose Fault-Aware Symbol Transformation (FAST) coupled with bit-cell natural redundancy reprogramming: FAST dynamically corrects computation errors for ±1 weights, while redundancy reprogramming restores stuck-at-zero weights. Evaluated on BitNet b1.58 and Wikitext, our approach reduces perplexity by 35% under SAFs, incurring only <3% energy overhead, <7% latency penalty, and <1% area cost. This work introduces the first tightly integrated framework combining symbolic transformation with hardware-level redundancy reprogramming, establishing a novel, high-accuracy, low-overhead fault-tolerance paradigm for TCiM accelerators.

Technology Category

Machine Learning: Hardware-aware MLConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSecurity and Privacy: Large-scale security measurementsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Ternary large language models (LLMs), which utilize ternary precision weights and 8-bit activations, have demonstrated competitive performance while significantly reducing the high computational and memory requirements of full-precision LLMs. The energy efficiency and performance of Ternary LLMs can be further improved by deploying them on ternary computing-in-memory (TCiM) accelerators, thereby alleviating the von-Neumann bottleneck. However, TCiM accelerators are prone to memory stuck-at faults (SAFs) leading to degradation in the model accuracy. This is particularly severe for LLMs due to their low weight sparsity. To boost the SAF tolerance of TCiM accelerators, we propose ReTern that is based on (i) fault-aware sign transformations (FAST) and (ii) TCiM bit-cell reprogramming exploiting their natural redundancy. The key idea is to utilize FAST to minimize computations errors due to SAFs in +1/-1 weights, while the natural bit-cell redundancy is exploited to target SAFs in 0 weights (zero-fix). Our experiments on BitNet b1.58 700M and 3B ternary LLMs show that our technique furnishes significant fault tolerance, notably 35% reduction in perplexity on the Wikitext dataset in the presence of faults. These benefits come at the cost of<3%,<7%, and<1% energy, latency and area overheads respectively.
Problem

Research questions and friction points this paper is trying to address.

Enhancing fault tolerance in ternary LLMs on TCiM accelerators
Mitigating memory stuck-at faults in low sparsity LLMs
Reducing computation errors and improving energy efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fault-aware sign transformations for error minimization
TCiM bit-cell reprogramming using natural redundancy
Energy-efficient ternary LLM deployment on TCiM accelerators
🔎 Similar Papers