Score
Designs and builds mechanisms that use one model’s outputs, feedback, or editable guidance to supervise, correct, or iteratively evolve the behavior of another model. This includes creating processes for generating corrective notes, hints, or policy edits to transfer capabilities and improve a weaker solver’s performance without directly updating its weights.
This study investigates the evolutionary mechanisms of self-evolving agent skills under multi-round feedback, focusing on how feedback type—success, failure, or both—affects skill refinement and whether test-time computation can reproduce the observed evolutionary gains. By constructing a controlled evaluation framework across five benchmarks and three models, the authors employ a multi-round feedback design with fixed executors, optimizers, and validation rules, complemented by byte-level difference detection and a validation-based selection mechanism. Their findings reveal that skill self-evolution is fundamentally a sparse search process critically dependent on validation filtering. Across 14 experimental settings, 11 successfully selected evolved skills, with 9 demonstrating improved test performance; notably, all effective evolutions required feedback incorporating failure trajectories. Moreover, test-time scaling with GPT-5.5 fails to fully recover these gains, highlighting a misalignment between validation criteria and downstream evaluation preferences.
Existing adaptive agent-based regulatory simulations lack mechanisms to systematically integrate diagnostic feedback into policy controllers, resulting in delayed and opaque policy adjustments. This work proposes a lightweight machine-guided policy revision layer that represents policies as defeasible rules and combines symbolic control, defeasible logic, and policy prioritization to operationalize contestability at the controller level, thereby endowing policy decisions with explainability, contestability, and dynamic revisability. In an emissions regulation agent-based model, the approach significantly reduces problem recurrence under scenarios where the VPVA mechanism fails due to excessive conservatism, while effectively maintaining key performance indicators such as violation rates, overshoot, and volatility.
研究解决了小模型在特定任务上表现不佳的问题,通过结合演化系统和在线策略修正方法,使小模型能够更好地适应并执行任务。
This work addresses the limitations of traditional self-evolving agents, which rely on predefined optimization pipelines that constrain large language models’ ability to autonomously plan improvement trajectories. The authors propose an Open-Ended Optimization (OEO) framework that, under fixed objectives, interaction boundaries, and resource budgets, enables optimizers to dynamically construct their own refinement processes online, redefining rigid pipelines as optional capability scaffolds. For the first time, they demonstrate that a state-of-the-art model (GPT-5.5) can autonomously accomplish complex optimization tasks without any preset workflow. Across 14 comparisons spanning eight benchmark–target model configurations, OEO achieves 12 wins, 1 tie, and 1 loss, while using only a median of 34.3% of the interaction token budget required by SkillOpt, substantially improving both efficiency and performance.
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
Existing agent evolvers rely on manually designed, fixed search loops and lack autonomous decision-making capabilities. This work proposes a meta-evolution framework that reformulates the optimization process itself as a learnable skill, enabling the evolver to autonomously determine testing, evaluation, and termination strategies. Specifically, the proposed method iteratively refines editable evolution skills by scoring newly generated agents, thereby achieving fully automated and adaptive pipelines. Experimental results demonstrate that this framework yields an average improvement of 13.6 points on primary metrics across multiple benchmarks, significantly outperforming handcrafted approaches. Furthermore, the acquired evolution skills exhibit strong transferability across diverse environments.
This study addresses the high computational costs of test-time training for large language models and the credit assignment challenges arising from coupled policy and implementation. We propose Guidance-TTT, a framework that introduces a novel guidance-execution decoupled architecture. Specifically, it freezes a large model to handle code implementation while exclusively training a small model to optimize high-level decision-making policies, thereby restricting test-time learning to short-horizon decision sequences. Furthermore, we design an adaptive group-relative reinforcement learning objective that leverages verifier feedback to enable online policy updates. Experimental results demonstrate that, in offline settings, this framework outperforms state-of-the-art methods across four major domains, including combinatorial optimization and machine learning, while significantly reducing computational overhead and enhancing exploration efficiency.
This study addresses the limitation of scalar feedback in automated large language model (LLM) research, where it fails to reveal behavioral conflicts that cause performance stagnation. To overcome this, we propose a competitive behavior feedback mechanism that identifies such conflicts and designs probe metrics to guide code agents in optimizing models. Furthermore, we introduce a reusable ConflictGuide-Skill that integrates literature taxonomy with model evidence to conduct a two-stage evolutionary search for conflict mitigation. Experimental results demonstrate that our approach reduces task error rates by 28% and conflict-related error rates by 14% across five model families, significantly outperforming baseline methods that rely solely on scalar feedback.
This study addresses the failure of scientific discovery caused by missing mechanisms under a fixed hypothesis space. We propose an experimental model class revision method that unifies mechanism expressibility and evidence collection into a single sequential decision problem. This approach jointly generates structural edits and diagnostic experiments, triggering theory revision through sequential evidence. To solve this formulation, we introduce a class-level distinguishability objective, anytime-valid sequential evidence, and reinforcement learning-based policies. Evaluated across 400 environments, our method achieves an exact recovery rate of 89.5%, significantly outperforming existing baselines while demonstrating strong transferability. These results establish a new paradigm for automated scientific discovery.
This study addresses the prohibitive costs and inefficiencies associated with model training and validation during AI agent evolution. To this end, we propose an idea-level critic model designed to predict the efficacy of modification proposals and filter high-potential strategies, thereby substituting expensive empirical validation. Methodologically, this specialized model is trained by integrating supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO) reinforcement learning, enabling optimized allocation of validation resources. Experimental results demonstrate that the proposed critic model outperforms existing frontier large language models. Across diverse tasks—including reasoning evolution, continual learning, and policy training—it significantly enhances both evolutionary efficiency under constrained budgets and the quality of final solutions.