RoSA: Rotational Sparse Adaptation for Memory-Efficient Fine-Tuning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the prohibitive memory overhead of fine-tuning large language models by proposing Rotational Sparse Adaptation. The method freezes lower-layer parameters and employs an inter-layer rotation mechanism for block-wise training, which integrates orthogonally with parameter-efficient fine-tuning (PEFT) techniques. Furthermore, it incorporates activation caching and sparse optimizer strategies to substantially reduce both optimizer state storage and backpropagation computational costs. Experimental results demonstrate that this framework significantly decreases peak memory consumption across various mainstream architectures while maintaining strong and stable fine-tuning performance. Consequently, it provides an efficient and generalizable solution for adapting large models under low-resource constraints.
📝 Abstract
Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
Problem

Research questions and friction points this paper is trying to address.

Parameter-efficient fine-tuning
Memory efficiency
Foundation models
Large language models
Optimizer-state memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rotational Sparse Adaptation
Parameter-Efficient Fine-Tuning
Memory-Efficient
Layer Freezing
Sparse Optimization
🔎 Similar Papers
No similar papers found.
M
Muhammad Azeem Lodhi
Saarland University, Saarbrücken, Germany
C
Chao Zhou
CISPA Helmholtz Center for Information Security, Saarbrücken, Germany
Rebekka Burkholz
Rebekka Burkholz
CISPA Helmholtz Center for Information Security
Machine LearningDeep Learning EfficiencyComplex NetworksCascadesGene Regulation