A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models

📅 2025-01-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the inefficiency, deployment overhead, and poor generalization of large language models (LLMs) in multi-task settings—including multilingual support, question answering (QA), and structured output generation—this paper proposes a multi-encoder–frozen-decoder architecture: the entire decoder is frozen, while only lightweight, task-specific encoders and adapter modules are fine-tuned. This work presents the first systematic empirical validation of decoder freezing under realistic multi-task and multilingual configurations, overcoming the limitation of prior frozen-parameter approaches—which were largely confined to single-task or representation-learning scenarios. Built upon the AlexaTM model and the PEFT paradigm, our method substantially reduces training cost and mitigates catastrophic forgetting. It maintains state-of-the-art (SOTA) performance on natural language generation (NLG) tasks, and—remarkably—outperforms full-parameter fine-tuning baselines on QA and structured generation tasks, achieving an unprecedented balance between deployment efficiency and cross-task generalization.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsMultiagent Systems: Multiagent Learning

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Among parameter-efficient fine-tuning methods, freezing has emerged as a popular strategy for speeding up training, reducing catastrophic forgetting, and improving downstream performance. We investigate the impact of freezing the decoder in a multi-task setup comprising diverse natural language tasks, aiming to reduce deployment overhead and enhance portability to novel tasks. Our experiments, conducted by fine-tuning both individual and multi-task setups on the AlexaTM model, reveal that freezing decoders is highly effective for tasks with natural language outputs and mitigates catastrophic forgetting in multilingual tasks. However, we find that pairing frozen decoders with a larger model can effectively maintain or even enhance performance in structured and QA tasks, making it a viable strategy for a broader range of task types.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Multi-task Learning
Structured Data Processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Angular Learning
Frozen Decoder
Enhanced Multitask Performance
🔎 Similar Papers
No similar papers found.
K
Kaustubh D. Dhole
Department of Computer Science, Emory University, Atlanta, USA