When Fine-Tuning Fails: Lessons from MS MARCO Passage Ranking

📅 2025-06-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study identifies an anomalous performance degradation—where fine-tuning strong pretrained Transformer models (e.g., BERT, DeBERTa) on the MS MARCO passage ranking task reduces MRR@10 below the base model’s score (0.3026)—across all five fine-tuning strategies, including full-parameter tuning and parameter-efficient methods like LoRA. Using UMAP-based embedding visualization, training dynamics analysis, and efficiency evaluation, we demonstrate that fine-tuning disrupts the optimal semantic embedding structure acquired during large-scale pretraining, inducing embedding space flattening. This challenges the conventional “fine-tuning always improves performance” paradigm in transfer learning—particularly on saturated benchmarks—and provides the first systematic evidence that excessive fine-tuning can impair retrieval ranking capability. Our findings suggest that architectural innovations or more robust adaptation mechanisms—not merely improved fine-tuning protocols—are critical to overcoming current limitations in dense retrieval.

Technology Category

Machine Learning: Learning Preferences or RankingsNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.Search and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
This paper investigates the counterintuitive phenomenon where fine-tuning pre-trained transformer models degrades performance on the MS MARCO passage ranking task. Through comprehensive experiments involving five model variants-including full parameter fine-tuning and parameter efficient LoRA adaptations-we demonstrate that all fine-tuning approaches underperform the base sentence-transformers/all- MiniLM-L6-v2 model (MRR@10: 0.3026). Our analysis reveals that fine-tuning disrupts the optimal embedding space structure learned during the base model's extensive pre-training on 1 billion sentence pairs, including 9.1 million MS MARCO samples. UMAP visualizations show progressive embedding space flattening, while training dynamics analysis and computational efficiency metrics further support our findings. These results challenge conventional wisdom about transfer learning effectiveness on saturated benchmarks and suggest architectural innovations may be necessary for meaningful improvements.
Problem

Research questions and friction points this paper is trying to address.

Fine-tuning degrades MS MARCO passage ranking performance
Fine-tuning disrupts optimal pre-trained embedding space structure
Challenges transfer learning effectiveness on saturated benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuning degrades MS MARCO ranking performance
LoRA adaptations underperform base model
Embedding space structure disrupted by fine-tuning
🔎 Similar Papers
No similar papers found.
M
Manu Pande
Department of IT, IIIT Allahabad, Prayagraj, India
S
Shahil Kumar
Department of IT, IIIT Allahabad, Prayagraj, India
A
Anay Yatin Damle
Department of IT, IIIT Allahabad, Prayagraj, India