Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the slow or unstable convergence caused by fixed step sizes in large model fine-tuning by proposing ZFO, a novel framework that pioneers a hybrid zeroth- and first-order optimization paradigm. The method decouples the update direction from the step size: it leverages first-order gradients to determine the update direction while requiring only two additional zeroth-order evaluations to estimate local curvature and construct a surrogate objective model, thereby enabling the adaptive selection of curvature-aware step sizes. Rigorous theoretical convergence guarantees are provided for the proposed approach. Extensive experiments across various large language model tasks validate its effectiveness, demonstrating that ZFO significantly outperforms fixed-step-size baselines and substantially improves final performance.
📝 Abstract
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.
Problem

Research questions and friction points this paper is trying to address.

Step-size selection
Large-scale neural network optimization
LLM fine-tuning
Convergence stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-and-First-Order Optimization
Step-size Selection
LLM Fine-Tuning
Curvature-aware Step
Decoupled Direction and Step