🤖 AI Summary
This work addresses the challenge of balancing computational efficiency and model performance in parameter-efficient fine-tuning of large language models by proposing TLoRA+, a novel fine-tuning method. Building upon the low-rank adaptation (LoRA) framework, TLoRA+ innovatively integrates a dedicated optimizer directly into the weight matrices of pretrained models. This design enhances fine-tuning effectiveness without introducing additional inference latency or substantially increasing computational overhead. Extensive experiments across multiple mainstream large language model architectures and the GLUE benchmark demonstrate that TLoRA+ consistently outperforms existing fine-tuning strategies, achieving superior performance and robustness while preserving the efficiency advantages of low-rank adaptation.
📝 Abstract
Fine-tuning large language models (LLMs) aims to adapt pre-trained models to specific tasks using relatively small and domain-specific datasets. Among Parameter-Efficient Fine-Tuning (PEFT) methods, Low-Rank Adaptation (LoRA) stands out by matching the performance of full fine-tuning while avoiding additional inference latency. In this paper, we propose a novel PEFT method that incorporates the TLoRA+ optimizer into the weight matrices of pre-trained models. The proposed approach not only preserves the efficiency of low-rank adaptation but also further enhances performance without significantly increasing computational cost. We conduct experiments on the GLUE benchmark across diverse model architectures. Numerical experiments consistently demonstrate the effectiveness and robustness of our proposed method.