Score
Design and implement training algorithms, inference pipelines, and evaluation/release workflows that use closed-loop feedback — such as inference-time critiques, critic scores, step-wise rewards, or post-training signals — to condition updates, trigger regenerations, or apply single-step RL refinements so models can self-correct during sampling. This includes building critic models and feedback-aware objectives, computing step-wise rewards or accuracy metrics, and instrumenting post-training feedback loops and conditional update logic for evaluation and release.
This work addresses the high cost, poor scalability, and diminishing effectiveness of human-supervised approaches for improving large language models, especially as model capabilities approach human-level performance. To overcome these limitations, the paper proposes a closed-loop self-improvement framework that structures the self-enhancement process into four tightly coupled stages: data acquisition, selection, model optimization, and inference refinement. A key innovation is the introduction of an autonomous evaluation layer that coordinates and guides transitions across these stages. This framework offers the first systematic, lifecycle-oriented modeling of self-improvement, unifying critical components such as self-generated data, automated evaluation, iterative training, and inference-time optimization. By comprehensively mapping existing technical pathways and their limitations, the study lays the groundwork for realizing fully autonomous, self-evolving language models.
This work addresses the limitation of existing tool-calling evaluation methods, which predominantly rely on post-hoc analysis and thus cannot correct errors in real time during reasoning. To overcome this, the authors propose a dual-agent architecture featuring an independent reviewer agent that evaluates tool calls before execution, shifting the paradigm from passive correction to proactive intervention. The reviewer’s decisions are guided by a Helpfulness-Harmlessness metric that quantifies the trade-off between potential benefits and risks. Coupled with an inference-time feedback mechanism and GEPA-based automatic prompt optimization, this approach enhances system performance without requiring model retraining. Empirical results demonstrate accuracy improvements of 5.5% and 7.1% on BFCL and Tau2-Bench, respectively, with the o3-mini model achieving a benefit-to-risk ratio of 3:1; further gains of 1.5–2.8% are attributable to GEPA optimization.
This work addresses the optimal allocation of post-training compute resources for reinforcement learning under a fixed FLOP budget. It introduces the first accounting framework that explicitly decomposes post-training computation into rollout/search, policy updates, and reward model evaluation, systematically quantifying the trade-offs among model scale, search intensity, number of learning steps, and feedback quality. Using GRPO with LoRA fine-tuning on the Qwen2.5 model family and combining rule-based and PRM rewards, the authors conduct large-scale ablation studies under a unified compute budget. Their findings reveal that the optimal allocation is highly sensitive to model size, total budget, reward type, and evaluation objective; notably, larger models incur higher per-inference costs, yielding fewer updates or rollouts within the same FLOP budget, thereby uncovering nonlinear coupling in compute allocation.
Current research on AI self-improvement lacks a systematic distinction between types of improvement and the degree of human-AI closed-loop interaction, impeding a clear understanding of the boundaries and risks of recursive self-improvement. This work addresses this gap by analyzing 1,250 arXiv papers and proposing the first dual-axis taxonomy that integrates improvement objectives—encompassing deployment behavior, policy training, evaluator optimization, and scientific research workflows—with levels of closed-loop autonomy. The study highlights the central role of self-evaluation signals within validation hierarchies and demonstrates that the intensity of self-improvement critically depends on the validation level at which these signals operate. It identifies “research direction setting” as a key bottleneck requiring human intervention and underscores governance-level metrics for self-improvement as the most underexplored area in current scholarship.
This work addresses the limitation of existing code generation models that rely solely on correctness feedback while neglecting runtime efficiency, often yielding syntactically correct but suboptimal solutions. To overcome this, the authors propose RLPF, a reinforcement learning–based approach featuring a staged composite reward mechanism: during early training stages, it provides execution-progress feedback for incorrect code, and upon achieving correctness, shifts focus to relative performance improvement using expert-written code as an efficiency benchmark. Fine-tuning the Qwen3-32B model with this framework significantly enhances both correctness and efficiency—on PerfCodeBench, the proportion of correct and efficient solutions rises from 11.1% to 54.6%, and relative efficiency improves from 8.1% to 38.6%. The method’s generalization capability is further validated on EffiBench-X, marking a departure from conventional correctness-only training paradigms.
This study investigates the evolutionary mechanisms of self-evolving agent skills under multi-round feedback, focusing on how feedback type—success, failure, or both—affects skill refinement and whether test-time computation can reproduce the observed evolutionary gains. By constructing a controlled evaluation framework across five benchmarks and three models, the authors employ a multi-round feedback design with fixed executors, optimizers, and validation rules, complemented by byte-level difference detection and a validation-based selection mechanism. Their findings reveal that skill self-evolution is fundamentally a sparse search process critically dependent on validation filtering. Across 14 experimental settings, 11 successfully selected evolved skills, with 9 demonstrating improved test performance; notably, all effective evolutions required feedback incorporating failure trajectories. Moreover, test-time scaling with GPT-5.5 fails to fully recover these gains, highlighting a misalignment between validation criteria and downstream evaluation preferences.
Existing reinforcement learning approaches struggle to disentangle initial code generation quality from iterative self-repair capabilities in multi-turn code generation and often overlook intermediate execution signals. This work proposes TaPR, a framework that introduces a unified multi-turn interaction protocol to transform execution feedback into fine-grained test-passing-rate rewards, enabling the first decoupled evaluation of initial generation and self-repair performance. TaPR incorporates a reward decomposition mechanism and a turn-aware evaluation protocol, optimized through a dense reward strategy based on test pass rates. Experiments demonstrate that TaPR improves the three-turn pass rate (Pass@3) by 2.44 percentage points on LiveCodeBench and boosts accuracy from 30.25% to 33.56% on the 7B/8B model subset, significantly outperforming baseline methods.
This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.
This work addresses the critical need for reliable step-level confidence estimation in large language model (LLM) agents, where single-step errors can lead to severe task failures. The authors propose a self-evolving critic framework that, for the first time, incorporates feedback on the consequences of executed actions into confidence assessment—without requiring ground-truth labels or additional training. By retrospectively evaluating outcomes to generate pseudo-labels, the method constructs and retrieves an experience memory bank, dynamically calibrating confidence through the integration of historical success and failure evidence when similar steps recur. Evaluated across three agent benchmarks and three backbone models, the approach significantly outperforms existing training-free baselines, achieving up to a 54% reduction in Expected Calibration Error (ECE) and setting new state-of-the-art results in both Brier score and AUC-based ranking performance.