🤖 AI Summary
This study addresses catastrophic forgetting and local overfitting arising from prompt optimization in the continual learning of large language models. To this end, we propose EFRE, a method that transcends the single-prompt paradigm by constructing a dynamically evolving function library. Through mechanisms that either update existing functions or instantiate new ones, EFRE adapts to novel tasks while effectively balancing the retention of prior knowledge with the acquisition of new capabilities. Experimental results demonstrate that EFRE improves average performance by 7.5% over GRPO across a three-task stream, with only a marginal 1.56% degradation on the FinQA task. These findings indicate significant superiority over existing baselines, establishing EFRE as an efficient solution for the continual learning of agent-based systems.
📝 Abstract
Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids costly parameter updates while achieving competitive or even superior performance to reinforcement learning methods such as GRPO on individual knowledge-intensive and reasoning tasks. This raises a natural question: \textit{Can prompt optimization, as an efficient adaptation approach, be directly applied to continual learning?} Our analysis shows that, under sequential task adaptation, it suffers from catastrophic forgetting, while optimized prompts accumulate rules that overfit to local task distributions. To address these limitations, we propose \emph{Evolving Functional REpertoires} (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones. On a three-task continual-learning stream, EFRE achieves a final average performance 7.50 percentage points higher than GRPO. Moreover, after adaptation to the Bio task, its performance on FinQA decreases by only 1.56 percentage points, compared with 25.10 percentage points for the base prompt optimization method. We further instantiate EFRE in a minimal agent system and observe consistent improvements across different backbone models. Overall, these results demonstrate EFRE's strong performance in continual learning for large language models and highlight its substantial potential for continual learning in advanced agent systems.