fine-tune encoder classifiers

Designs, implements, and evaluates procedures to adapt pretrained encoder-based models and classifiers—including encoder-only language models, acoustic encoders, multilingual and multimodal encoders—by fine-tuning weights, representation layers, and hyperparameters (e.g., learning rate and parameter subsets). This work produces task-specific encoders and classifier heads through post-training fine-tuning and optimization of trade-offs such as latency, cost, and predictive performance.

fine-tuneencoderclassifiers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.98
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work investigates the hierarchical representation evolution of pretrained Transformers (e.g., BERT) during fine-tuning—specifically, how models balance preservation of general-purpose features against acquisition of task-specific ones. We propose an activation interpretability framework based on sparse autoencoders (SAEs), integrating inter-layer similarity metrics, token-level activation visualization, and cross-dataset comparative experiments. Our analysis systematically reveals, for the first time, a universal layer-wise functional specialization pattern: early layers retain generic linguistic features, middle layers exhibit progressive transitional behavior, and deeper layers specialize in task-adaptive representations. This finding provides both theoretical grounding and empirical evidence for controllable fine-tuning, model compression, and interpretable AI.

Analyzing fine-tuning mechanisms in transformersExploring feature adaptation across BERT layersUnderstanding task-specific vs. general representations

This study addresses the insufficient exploration of relationships and trade-offs among different paradigms for cross-task adaptation of large language models by constructing the first unified analytical framework. Methodologically, it systematically integrates parameter-efficient fine-tuning, in-context learning, and embedding injection techniques, establishing a comprehensive taxonomy along the dimensions of weights, prompts, and embeddings. Employing a systematic review methodology, the work thoroughly analyzes the strengths, limitations, and intrinsic connections of each paradigm. The findings reveal the evolutionary logic underlying these three paradigms, yielding a holistic taxonomic landscape that clarifies their core advantages and constraints while identifying key open problems. Ultimately, this research provides strategic guidance for future investigations into the efficient adaptation of large models.

Embedding-Based AdaptationIn-Context LearningLarge Language Models

Scalable Fine-tuning from Multiple Data Sources:A First-Order Approximation Approach

Sep 28, 2024
DL
Dongyue Li
🏛️ Northeastern University | University of Michigan

This work addresses the problem of efficiently selecting an optimal subset of auxiliary tasks from multiple sources to enhance fine-tuning performance on a target task. We propose a lightweight, training-free subset selection method grounded in meta-initialization. Our core contribution is a first-order gradient approximation framework that leverages first-order Taylor expansion and gradient sensitivity analysis to estimate fine-tuning loss for arbitrary task subsets—entirely on CPU, within seconds. Unlike conventional enumeration or reinforcement learning–based approaches, our method eliminates the need for repeated fine-tuning, achieving a 30× speedup over baselines with only 1% estimation error. Empirically, on instruction tuning and chain-of-thought tuning benchmarks, subsets selected by our method yield up to a 3.8% absolute improvement in downstream task performance, significantly outperforming existing subset selection techniques.

Accurately estimating fine-tuning performance via gradient approximationAvoiding repeated training in subset selection for efficiencyOptimally selecting beneficial auxiliary tasks for LM fine-tuning

This work addresses the lack of systematic guidance in existing parameter-efficient fine-tuning (PEFT) methods regarding which layers of large language models should be adapted, a gap that complicates the trade-off among performance, training cost, and inference latency. The authors propose a unified projected residual perspective, revealing through local quadratic approximation that inter-layer adaptability is governed by three factors: residual norm, activation energy, and layer coupling. Leveraging this insight, they develop a reusable diagnostic tool—Layer Card—that enables goal-directed, flexible layer selection. This framework provides the first theoretical foundation for layer selection in PEFT, demonstrating on Qwen3-8B that fine-tuning only a subset of layers can closely match the performance of full-layer LoRA while substantially reducing both training overhead and the number of active adapters during inference.

fine-tuning costinference latencylarge language models

How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario

Nov 27, 2024
SW
Shih-Heng Wang
🏛️ National Taiwan University | Carnegie Mellon University | The University of Texas at Austin | FAIR

To address the limited cross-lingual transferability and severe domain mismatch of self-supervised learning (SSL) pre-trained models in automatic speech recognition (ASR) for low-resource languages, this paper proposes a lightweight adapter method with *intermediate warm-start*. Under frozen SSL backbone constraints, only 1–5% of parameters are fine-tuned. A two-stage progressive adaptation jointly optimizes adapter architecture and downstream model initialization. The novel intermediate warm-start mechanism mitigates speech feature distribution shift, substantially improving generalization to unseen languages. Evaluated on the ML-SUPERB benchmark, our approach achieves up to 28% relative reduction in character/phone error rates over standard efficient fine-tuning, significantly alleviating the bottleneck in low-resource cross-lingual ASR adaptation.

Automatic Speech Recognition (ASR)Low-Resource LanguagesSpeech Self-Supervised Learning (SSL)

Latest Papers

What's happening recently
View more

This work addresses the challenge of selecting effective fine-tuning strategies for encoder-decoder pre-trained language models in generation and question-answering tasks. It proposes the Match Task to Objective (MTO) framework, which establishes the first systematic alignment mechanism between downstream tasks and pre-training objectives. MTO automatically constructs training data and prompt templates that are consistent with the original pre-training objective and extends this alignment to soft prompt tuning, thereby enabling precise task–objective matching. Experimental results demonstrate that MTO achieves over 120% performance improvement under few-shot settings compared to existing methods, significantly outperforms strong baselines in full-data scenarios, and substantially enhances the effectiveness of prompt tuning.

commonsense reasoningencoder-decoder modelspre-training objectives

Pretraining for low-resource languages is often constrained by data scarcity, necessitating repeated use of limited corpora and thereby compromising model generalization. This work systematically compares two strategies to mitigate this issue: mixed training with high-resource auxiliary languages and extensive hyperparameter tuning. Through large-scale hyperparameter search (approximately 1,000 runs) combined with μP-based transfer tuning, we demonstrate that mixed training substantially outperforms fine-tuned monolingual baselines—a gap that widens with increasing model scale from 150M to 1.43B parameters. Mixed training yields validation loss improvements equivalent to a 2–3× increase in target-language data and downstream accuracy gains equivalent to a 2–13× data boost. Moreover, we show that validation loss measured solely on the target language systematically underestimates the cross-lingual knowledge benefits conferred by mixed training.

bilingual pre-trainingdata-constrainedgeneralization

This study systematically investigates whether representation collapse during multi-stage post-training of large language models leads to degraded adaptability, weakened out-of-distribution generalization, and deteriorated calibration. To this end, the authors construct a comprehensive measurement framework encompassing hidden states, logits, token trajectories, and LoRA updates, thereby establishing— for the first time—the causal relationship between representation collapse and declines in model plasticity, generalization, and calibration. Building on these insights, they propose lightweight intervention strategies, including mixed-domain replay, feature refreshing, representation diversity regularization, and decorrelation of LoRA updates. These methods significantly enhance continual learning capability while preserving behavioral gains, effectively mitigating representation collapse.

adaptation plasticitylarge language modelsrepresentation collapse

This study addresses the inherent trade-off between task-specific performance gains and general capability degradation during large language model fine-tuning by proposing LS-LoRA. The method identifies inter-layer sensitivity variations and introduces input-output cosine similarity as a lightweight forward-pass proxy metric, effectively replacing computationally expensive empirical Fisher information calculations. Based on this metric, LS-LoRA selectively places LoRA adapters in low-sensitivity layers. Experimental results demonstrate that LS-LoRA substantially improves target performance on mathematical and coding tasks while largely preserving general capabilities such as commonsense reasoning. Overall, this work achieves an effective balance between task adaptation and capability preservation under parameter-efficient fine-tuning paradigms.

Capability retentionCatastrophic forgettingLarge language models

This work addresses the inefficiency of full-model fine-tuning in monolingual language model development for low-resource languages, which incurs high computational costs and fails to leverage modular adaptation effectively. To overcome this limitation, the authors propose an efficient transfer learning approach that integrates a target-language-specific tokenizer, freezes the corresponding embedding layer, and fine-tunes only the remaining model parameters—departing from conventional full-parameter fine-tuning paradigms. Evaluated on Scottish Gaelic, Irish, and Quechua (with only 8.5k training samples), the method consistently outperforms baseline approaches across masked language modeling, named entity recognition, and part-of-speech tagging tasks, demonstrating its effectiveness and generalizability under extremely low-resource conditions.

low-resource languagesmodel adaptationmonolingual adaptation

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
YF

Yanwei Fu

Fudan University
Computer visionmachine learningMultimedia
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
ST

Shengbang Tong

NYU Courant
AIComputer VisionDeep LearningRepresentation Learning
BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain