COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the disconnect between parameter updates and context optimization in the recursive self-improvement of large language models, where their interaction is typically overlooked. To this end, we propose COEVO, a framework that formally defines the parameter-context co-evolution problem for the first time. COEVO treats model parameters and learning contexts as a co-evolving entity that adapts synchronously within a shared feedback loop, leveraging policy entropy and prompt-conditioned attention as complementary signals to guide adaptive adjustments. Experimental results demonstrate that COEVO significantly outperforms reinforcement learning baselines with fixed contexts in task performance while substantially enhancing policy robustness against variations in system prompts. These findings validate the core value of context-adaptive optimization in advancing recursive self-improvement.
πŸ“ Abstract
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter--context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.
Problem

Research questions and friction points this paper is trying to address.

Recursive Self-Improvement
Large Language Models
Parameter-Context Co-evolution
Online Learning
Prompt Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
Co-Evolution
On-Policy Reinforcement Learning
Policy Entropy
Prompt-Conditioned Attention