closed-loop feedback training

Design and implement training algorithms, inference pipelines, and evaluation/release workflows that use closed-loop feedback — such as inference-time critiques, critic scores, step-wise rewards, or post-training signals — to condition updates, trigger regenerations, or apply single-step RL refinements so models can self-correct during sampling. This includes building critic models and feedback-aware objectives, computing step-wise rewards or accuracy metrics, and instrumenting post-training feedback loops and conditional update logic for evaluation and release.

closed-loopfeedbacktraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing tool-calling evaluation methods, which predominantly rely on post-hoc analysis and thus cannot correct errors in real time during reasoning. To overcome this, the authors propose a dual-agent architecture featuring an independent reviewer agent that evaluates tool calls before execution, shifting the paradigm from passive correction to proactive intervention. The reviewer’s decisions are guided by a Helpfulness-Harmlessness metric that quantifies the trade-off between potential benefits and risks. Coupled with an inference-time feedback mechanism and GEPA-based automatic prompt optimization, this approach enhances system performance without requiring model retraining. Empirical results demonstrate accuracy improvements of 5.5% and 7.1% on BFCL and Tau2-Bench, respectively, with the o3-mini model achieving a benefit-to-risk ratio of 3:1; further gains of 1.5–2.8% are attributable to GEPA optimization.

execution loopinference-time feedbackpost-hoc evaluation

This work addresses the optimal allocation of post-training compute resources for reinforcement learning under a fixed FLOP budget. It introduces the first accounting framework that explicitly decomposes post-training computation into rollout/search, policy updates, and reward model evaluation, systematically quantifying the trade-offs among model scale, search intensity, number of learning steps, and feedback quality. Using GRPO with LoRA fine-tuning on the Qwen2.5 model family and combining rule-based and PRM rewards, the authors conduct large-scale ablation studies under a unified compute budget. Their findings reveal that the optimal allocation is highly sensitive to model size, total budget, reward type, and evaluation objective; notably, larger models incur higher per-inference costs, yielding fewer updates or rollouts within the same FLOP budget, thereby uncovering nonlinear coupling in compute allocation.

compute allocationFLOP budgetfoundation models

Current research on AI self-improvement lacks a systematic distinction between types of improvement and the degree of human-AI closed-loop interaction, impeding a clear understanding of the boundaries and risks of recursive self-improvement. This work addresses this gap by analyzing 1,250 arXiv papers and proposing the first dual-axis taxonomy that integrates improvement objectives—encompassing deployment behavior, policy training, evaluator optimization, and scientific research workflows—with levels of closed-loop autonomy. The study highlights the central role of self-evaluation signals within validation hierarchies and demonstrates that the intensity of self-improvement critically depends on the validation level at which these signals operate. It identifies “research direction setting” as a key bottleneck requiring human intervention and underscores governance-level metrics for self-improvement as the most underexplored area in current scholarship.

AI SafetyGovernanceLoop Closure

This work addresses the limitation of existing code generation models that rely solely on correctness feedback while neglecting runtime efficiency, often yielding syntactically correct but suboptimal solutions. To overcome this, the authors propose RLPF, a reinforcement learning–based approach featuring a staged composite reward mechanism: during early training stages, it provides execution-progress feedback for incorrect code, and upon achieving correctness, shifts focus to relative performance improvement using expert-written code as an efficiency benchmark. Fine-tuning the Qwen3-32B model with this framework significantly enhances both correctness and efficiency—on PerfCodeBench, the proportion of correct and efficient solutions rises from 11.1% to 54.6%, and relative efficiency improves from 8.1% to 38.6%. The method’s generalization capability is further validated on EffiBench-X, marking a departure from conventional correctness-only training paradigms.

code generationexecution feedbackprogram optimization

This study investigates the evolutionary mechanisms of self-evolving agent skills under multi-round feedback, focusing on how feedback type—success, failure, or both—affects skill refinement and whether test-time computation can reproduce the observed evolutionary gains. By constructing a controlled evaluation framework across five benchmarks and three models, the authors employ a multi-round feedback design with fixed executors, optimizers, and validation rules, complemented by byte-level difference detection and a validation-based selection mechanism. Their findings reveal that skill self-evolution is fundamentally a sparse search process critically dependent on validation filtering. Across 14 experimental settings, 11 successfully selected evolved skills, with 9 demonstrating improved test performance; notably, all effective evolutions required feedback incorporating failure trajectories. Moreover, test-time scaling with GPT-5.5 fails to fully recover these gains, highlighting a misalignment between validation criteria and downstream evaluation preferences.

feedback dynamicsself-evolving agentsskill revision

Latest Papers

What's happening recently
View more

Existing reinforcement learning approaches struggle to disentangle initial code generation quality from iterative self-repair capabilities in multi-turn code generation and often overlook intermediate execution signals. This work proposes TaPR, a framework that introduces a unified multi-turn interaction protocol to transform execution feedback into fine-grained test-passing-rate rewards, enabling the first decoupled evaluation of initial generation and self-repair performance. TaPR incorporates a reward decomposition mechanism and a turn-aware evaluation protocol, optimized through a dense reward strategy based on test pass rates. Experiments demonstrate that TaPR improves the three-turn pass rate (Pass@3) by 2.44 percentage points on LiveCodeBench and boosts accuracy from 30.25% to 33.56% on the 7B/8B model subset, significantly outperforming baseline methods.

code generationexecution feedbackmulti-turn interaction

This work addresses the semantic gap between evaluation metrics and training data in large model pretraining, which hinders precise diagnosis and remediation of capability deficiencies. The authors propose “capability slices” as fundamental units aligning evaluation and data, establishing a bidirectional classification framework that links evaluation tasks with non-instructional training data through explicit mapping rules. This enables a closed-loop pipeline from evaluation failures to targeted data interventions. For the first time, the approach supports auditable and systematic reasoning that translates evaluation signals into data corrections, moving beyond intuition-driven tuning paradigms. Experiments demonstrate its efficacy in both directions: repairing specific training loss components restores BBH performance to 66.44, while targeted data sampling boosts AIME2025/2026 Pass@128 from 6.67/0.00 to 26.67.

capability slicedata-evaluation gapevaluation-to-data inference

This work addresses the critical need for reliable step-level confidence estimation in large language model (LLM) agents, where single-step errors can lead to severe task failures. The authors propose a self-evolving critic framework that, for the first time, incorporates feedback on the consequences of executed actions into confidence assessment—without requiring ground-truth labels or additional training. By retrospectively evaluating outcomes to generate pseudo-labels, the method constructs and retrieves an experience memory bank, dynamically calibrating confidence through the integration of historical success and failure evidence when similar steps recur. Evaluated across three agent benchmarks and three backbone models, the approach significantly outperforms existing training-free baselines, achieving up to a 54% reduction in Expected Calibration Error (ECE) and setting new state-of-the-art results in both Brier score and AUC-based ranking performance.

action productivitycalibrated probabilityexecution consequences

Hot Scholars

AA

Alexandre Alahi

Professor, EPFL
Computer VisionTransportationAutonomous drivingIntelligent Transportation Systems
WL

Wuyang Li

EPFL
video generative models2D/3D visual perceptionmeta-optics
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
ZL

Zhe Liu

The University of Hong Kong
3D PerceptionEmbodied AI4D MLLM4D World Model