Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

📅 2026-08-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether gradient-free weight perturbation methods require perturbing all model parameters and what factors truly govern their effectiveness. Through controlled ablation experiments, the work systematically disentangles the effects of perturbation dimensionality, subspace choice, and perturbation norm, revealing for the first time that the perturbation norm—not the dimensionality or subspace—is the key determinant of performance. Empirical results demonstrate that once norms are matched, diverse subspaces—such as those derived from SVD bases or random frames—yield nearly identical outcomes. Remarkably, perturbing as few as 12–16 scalar parameters achieves 98.2% of the performance of full-parameter perturbation (an average gap of only 1.8 percentage points). This insight exhibits strong cross-model transferability and offers a novel perspective on efficient fine-tuning.
📝 Abstract
Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a handful of scalars. Gradient-free adaptation, which samples random weight perturbations and keeps the ones that score well, has not followed that trajectory and still perturbs every entry of the weight tensor. It is unknown whether that full-weight search is necessary, and more fundamentally which property of a perturbation makes it work at all, because existing methods vary the search space, the perturbation scale, and the aggregation together. We resolve this by intervening on one factor at a time inside a fixed pipeline, holding candidate scoring and voting constant while we vary the search dimension, the subspace that carries the perturbation, and its norm. Perturbing a frozen frame of 12 to 16 scalars stays 1.8 accuracy points behind full-weight search on average across 49 model-benchmark cells, trailing it in 36 of them. Neither the dimension nor the choice of basis explains that performance. A random frame whose Grassmann overlap with the SVD frame is at chance level performs identically once a single scale factor is matched, and at large scales the SVD directions collapse first. What survives is the perturbation norm, whose usable range closes within a factor of five across seven models and stays flat inside. The perturbation norm is therefore the one factor with a failure mode, and its safe region transfers across scale and family. The design question narrows from which subspace to perturb to how hard to shake.
Problem

Research questions and friction points this paper is trying to address.

gradient-free adaptation
weight perturbation
language models
perturbation norm
parameter-efficient
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient-free adaptation
weight perturbation
perturbation norm
parameter-efficient tuning
Grassmann overlap
🔎 Similar Papers
No similar papers found.
T
Taeyeong Kim
Chosun University
A
Ahhyun Kim
Chosun University
T
TaeHyeon Kim
Chosun University
U
Unggi Lee
Korea University Sejong Campus