AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of isotropic perturbations and the limited quality of subspaces in zeroth-order fine-tuning of large language models. To this end, it proposes an activation-aware subspace-based zeroth-order fine-tuning method that leverages forward activations to construct dynamic subspaces for guiding parameter updates. The approach innovatively decouples maintenance width from active width, continuously integrating local directional information to capture high-quality gradient structures. Furthermore, by incorporating a shared sampling direction strategy, it achieves efficient optimization without backpropagation. Experimental evaluations on OPT-series models demonstrate that the proposed method outperforms the strongest baseline by 1.26 to 2.85 percentage points on average, substantially enhancing fine-tuning efficiency.
📝 Abstract
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly informative directions. To make these evaluations more informative, existing ZO methods restrict perturbations to low-dimensional subspaces. Yet the quality of these subspaces is critical: overly compressed or poorly maintained spaces can miss useful update directions. To obtain a high-quality subspace for ZO updates, this paper proposes AIM-ZO, a ZO fine-tuning method based on Activation-Informed Subspace Maintenance. AIM-ZO uses forward activations as local directional information and continuously integrates them into a broad, evolving subspace over training. To access broader gradient-relevant structure while keeping individual perturbations low-dimensional, AIM-ZO activates only a smaller set of shared and sampled directions, decoupling the maintained width from the active width. We evaluate AIM-ZO across 5 LLMs and 11 downstream tasks under matched forward-evaluation budgets; its six-task average exceeds the strongest fully evaluated ZO baseline by 1.26 percentage points on OPT-2.7B and MeZO by 2.85 percentage points on OPT-30B. Our code is available at https://github.com/EkkoXy/AIM-ZO
Problem

Research questions and friction points this paper is trying to address.

Zeroth-order optimization
LLM fine-tuning
Subspace maintenance
Memory efficiency
Forward evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zeroth-Order Optimization
Activation-Informed Subspace
LLM Fine-Tuning
Subspace Maintenance
Memory Efficiency
💼 Related Jobs
No related jobs found.