🤖 AI Summary
This study addresses the excessive activation memory consumption encountered when adapting convolutional neural networks (CNNs) for edge devices by proposing Memory-Floor LoRA. This method reconstructs low-rank adapters under a memory-first principle, introducing an activation memory lower-bound criterion that transcends the conventional limitation of optimizing only trainable parameters. By integrating strategies such as freezing down-projections, training up-projections, applying evaluation-mode normalization, and minimizing backward computation dependencies, it substantially reduces activation memory requirements. Experiments on human activity recognition tasks demonstrate that the proposed approach decreases activation memory by 98.5%–98.7% and peak training memory by 94.9%–97.3%, while achieving performance comparable to or exceeding that of existing baselines.
📝 Abstract
On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants, enable efficient adaptation at the edge, the limiting resource for Convolutional Neural Network (CNN) adaptation is often not the number of trainable parameters but the activation state that must be retained until the backward pass. This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA. Instead of merely reducing trainable weights, we define an activation-memory-floor criterion: trainable backward computations must not depend on full-width layer inputs. The resulting adapter freezes the down-projection, trains a scale-matched up-projection, and combines eval-mode backbone normalization with activation-minimal backward rules, reducing saved state to the low-rank branch. Evaluated on three Human Activity Recognition (HAR) datasets and two CNN backbones under subject, body-location, and sensor-placement shifts, MemFLoRA reduces saved-activation memory by 98.5-98.7% and peak training-state memory by 94.9-97.3% relative to full fine-tuning, while matching or exceeding CNN PEFT baselines.