🤖 AI Summary
To address the inefficiency, deployment overhead, and poor generalization of large language models (LLMs) in multi-task settings—including multilingual support, question answering (QA), and structured output generation—this paper proposes a multi-encoder–frozen-decoder architecture: the entire decoder is frozen, while only lightweight, task-specific encoders and adapter modules are fine-tuned. This work presents the first systematic empirical validation of decoder freezing under realistic multi-task and multilingual configurations, overcoming the limitation of prior frozen-parameter approaches—which were largely confined to single-task or representation-learning scenarios. Built upon the AlexaTM model and the PEFT paradigm, our method substantially reduces training cost and mitigates catastrophic forgetting. It maintains state-of-the-art (SOTA) performance on natural language generation (NLG) tasks, and—remarkably—outperforms full-parameter fine-tuning baselines on QA and structured generation tasks, achieving an unprecedented balance between deployment efficiency and cross-task generalization.
📝 Abstract
Among parameter-efficient fine-tuning methods, freezing has emerged as a popular strategy for speeding up training, reducing catastrophic forgetting, and improving downstream performance. We investigate the impact of freezing the decoder in a multi-task setup comprising diverse natural language tasks, aiming to reduce deployment overhead and enhance portability to novel tasks. Our experiments, conducted by fine-tuning both individual and multi-task setups on the AlexaTM model, reveal that freezing decoders is highly effective for tasks with natural language outputs and mitigates catastrophic forgetting in multilingual tasks. However, we find that pairing frozen decoders with a larger model can effectively maintain or even enhance performance in structured and QA tasks, making it a viable strategy for a broader range of task types.