LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers

📅 2026-06-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of explicitly controlling overfitting during fine-tuning of pretrained Transformers by formulating it as a bilevel optimization regularized framework. It introduces, for the first time, a linear programming–driven local search mechanism that leverages validation gradients and training Hessian information from a warm-up phase to construct a validation-aware descent direction. This enables joint, task-adaptive optimization of both model parameters and regularization hyperparameters without requiring repeated full retraining. Experimental results demonstrate significant reductions in test perplexity on GPT-2 Small and WikiText-2, with particularly pronounced gains in settings prone to overfitting. The approach consistently yields stable improvements across diverse layer configurations and regularization settings.
📝 Abstract
This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting. The approach formulates transformer fine-tuning as a bilevel optimization-based regularization problem, in which model parameters and regularization hyperparameters are jointly updated. Information collected during initial warm-up iterations, including validation gradients and training Hessian information, is used to construct a local descent direction by solving an LP that minimizes a scaled directional derivative while preserving training optimality. This validation-aware descent direction enables focused local updates of both parameters and regularization hyperparameters, reducing overfitting without requiring repeated full retraining cycles. The resulting method, termed Linear Programming-based Fine-Tuning (LiFT) for transformers, differs from conventional fine-tuning by systematically identifying task-specific updates rather than relying on heuristic or grid-based hyperparameter selection. Experiments on GPT-2 Small fine-tuned on WikiText-2 demonstrate that LiFT enables effective adaptation through selective tuning of transformer blocks and regularization parameters, yielding consistent improvements in test perplexity across multiple layer configurations and regularization settings, with particularly pronounced gains in overfitting-prone scenarios. Beyond empirical performance, LiFT establishes a principled connection between transformer fine-tuning, bilevel optimization, local search, and regularization theory.
Problem

Research questions and friction points this paper is trying to address.

overfitting
transformer fine-tuning
regularization
bilevel optimization
local search
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Programming
Bilevel Optimization
Overfitting Control
Local Search
Transformer Fine-tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Abhishek Shukla
Department of Management Sciences, Indian Institute of Technology Kanpur, Kanpur-208016, Uttar Pradesh, India
A
Anikeit Khanna
Department of Civil Engineering, Indian Institute of Technology Kanpur, Kanpur-208016, Uttar Pradesh, India
A
Ankur Sinha
Operations and Decision Sciences, Indian Institute of Management Ahmedabad, Ahmedabad-380015, Gujarat, India
F
Faiz Hamid
Department of Management Sciences, Indian Institute of Technology Kanpur, Kanpur-208016, Uttar Pradesh, India