Robust Parameter-Efficient LLM Adaptation on Analog Hardware

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation in large language model fine-tuning caused by hardware noise and precision limitations in analog compute-in-memory systems, where full retraining remains prohibitively expensive. To this end, we propose a parameter-efficient adaptation method based on Low-Rank Adaptation (LoRA). By integrating input reshaping with update accumulation techniques, our approach achieves optimizer-agnostic robust adaptation under fixed pre-trained weights, effectively overcoming errors in forward and backward matrix-vector multiplications as well as the challenges of physical weight updates under limited conductance states. Experimental results demonstrate that the proposed method significantly improves the fine-tuning performance of Llama-series models in noisy environments, maintaining superior efficacy even under stringent conditions restricted to merely 20 conductance states.
📝 Abstract
Analog in-memory computing is a promising platform for on-device execution of large language models because it performs matrix--vector multiplications (MVMs) in memory and in parallel, reducing data movement. However, limited digital-to-analog converter precision, input noise, and finite conductance states can degrade model accuracy, while full-model retraining to address these effects can be costly. We develop an optimizer-agnostic, parameter-efficient adaptation method based on Low-Rank Adaptation (LoRA), keeping the pretrained weights stored on analog arrays fixed while training the LoRA weights to adapt to downstream tasks and hardware non-idealities. Reliable adaptation requires handling errors in both forward and backward MVMs and physical weight updates. We use input reshaping to reduce input-induced MVM errors and update accumulation to retain small updates before programming them to finite-state analog devices. Across Llama-3.2-1B-Instruct and Llama-3-8B with both Muon and AdamW, input reshaping improves analog LoRA fine-tuning under noisy MVM computation. Update accumulation separately preserves sub-threshold updates and substantially improves adaptation under finite-resolution programming, including configurations with as few as 20 conductance states. Additional experiments show consistent held-out negative log-likelihood improvements across noisy analog settings.
Problem

Research questions and friction points this paper is trying to address.

Analog in-memory computing
Large language models
Parameter-efficient adaptation
Hardware non-idealities
Low-Rank Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analog In-Memory Computing
Low-Rank Adaptation (LoRA)
Parameter-Efficient Fine-Tuning
Input Reshaping
Update Accumulation
🔎 Similar Papers
No similar papers found.