Reassessing Global Gradient-Norm Imbalance in BLIP Fine-Tuning Across Physical Domains

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究测试了不同物理领域下BLIP微调时视觉和语言路径梯度不平衡问题,通过调整学习率、阶段性冻结等方法减少梯度失衡,但发现梯度平衡程度与字幕生成性能无稳定关联。
📝 Abstract
Imbalanced gradient magnitudes between the visual and language pathways of a vision-language model are often treated as a defect to be corrected. We test that premise for one family of correction, deliberately excluding adaptive, signal-driven schemes (e.g. BalGrad, OGM, PMR, CGGM), which are a mechanistically distinct class outside this study's scope. Measuring the language-to-visual gradient-norm ratio, reported in parameter-normalised form, across nine fine-tuning conditions, three seeds, and three captioning datasets spanning distinct physical domain shifts -- underwater, aerial, radiological -- we find imbalance magnitude varies markedly across domains with no predictable ordering. A plain learning-rate reduction cuts imbalance substantially and lands within a few BLEU points of the best method on every dataset. Staged freezing reduces the ratio on every domain yet never ranks first; a schedule-only control isolates freezing as the cause on one dataset but not the other two. Forcing the two gradient groups to equal magnitude drives per-parameter imbalance close to zero on every domain, yet is both the best result in the study and the worst placement among full fine-tuning methods, on different datasets, with identical settings. Reductions in gradient-norm ratio do not consistently predict captioning performance across domains, and how a given level of balance is reached matters as much as the level itself. As a secondary finding, a commonly reused LoRA configuration applied to BLIP silently adapts zero visual parameters; correcting it improves BLEU-4 on all three datasets.
Problem

Research questions and friction points this paper is trying to address.

Gradient-Norm Imbalance
Vision-Language Model
Fine-Tuning
Physical Domains
Captioning Performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient-norm imbalance
learning-rate reduction
staged freezing
vision-language model fine-tuning
LoRA configuration
💼 Related Jobs
No related jobs found.
K
Kiran Naseer
Department of Computer Science, University of Gujrat, Pakistan
S
Samreen Azhar
Department of Computer Science, University of Gujrat, Pakistan
Dwarikanath Mahapatra
Dwarikanath Mahapatra
Khalifa University
AI in MedicineMedical Image SegmentationMedical Image RegistrationComputer VisionDeep