Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive costs and inefficiencies associated with model training and validation during AI agent evolution. To this end, we propose an idea-level critic model designed to predict the efficacy of modification proposals and filter high-potential strategies, thereby substituting expensive empirical validation. Methodologically, this specialized model is trained by integrating supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO) reinforcement learning, enabling optimized allocation of validation resources. Experimental results demonstrate that the proposed critic model outperforms existing frontier large language models. Across diverse tasks—including reasoning evolution, continual learning, and policy training—it significantly enhances both evolutionary efficiency under constrained budgets and the quality of final solutions.
📝 Abstract
As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.
Problem

Research questions and friction points this paper is trying to address.

Self-evolving agents
AI for ML
Verification efficiency
Idea-level critics
Verification budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Idea-Level Critics
Verification-Efficient Agents
GRPO
AI4ML
Learned Reward Models