cross-model supervision

Designs and builds mechanisms that use one model’s outputs, feedback, or editable guidance to supervise, correct, or iteratively evolve the behavior of another model. This includes creating processes for generating corrective notes, hints, or policy edits to transfer capabilities and improve a weaker solver’s performance without directly updating its weights.

cross-modelsupervision

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the evolutionary mechanisms of self-evolving agent skills under multi-round feedback, focusing on how feedback type—success, failure, or both—affects skill refinement and whether test-time computation can reproduce the observed evolutionary gains. By constructing a controlled evaluation framework across five benchmarks and three models, the authors employ a multi-round feedback design with fixed executors, optimizers, and validation rules, complemented by byte-level difference detection and a validation-based selection mechanism. Their findings reveal that skill self-evolution is fundamentally a sparse search process critically dependent on validation filtering. Across 14 experimental settings, 11 successfully selected evolved skills, with 9 demonstrating improved test performance; notably, all effective evolutions required feedback incorporating failure trajectories. Moreover, test-time scaling with GPT-5.5 fails to fully recover these gains, highlighting a misalignment between validation criteria and downstream evaluation preferences.

feedback dynamicsself-evolving agentsskill revision

Existing adaptive agent-based regulatory simulations lack mechanisms to systematically integrate diagnostic feedback into policy controllers, resulting in delayed and opaque policy adjustments. This work proposes a lightweight machine-guided policy revision layer that represents policies as defeasible rules and combines symbolic control, defeasible logic, and policy prioritization to operationalize contestability at the controller level, thereby endowing policy decisions with explainability, contestability, and dynamic revisability. In an emissions regulation agent-based model, the approach significantly reduces problem recurrence under scenarios where the VPVA mechanism fails due to excessive conservatism, while effectively maintaining key performance indicators such as violation rates, overshoot, and volatility.

adaptive simulationagent-based modelingcontestability

This work addresses the limitations of traditional self-evolving agents, which rely on predefined optimization pipelines that constrain large language models’ ability to autonomously plan improvement trajectories. The authors propose an Open-Ended Optimization (OEO) framework that, under fixed objectives, interaction boundaries, and resource budgets, enables optimizers to dynamically construct their own refinement processes online, redefining rigid pipelines as optional capability scaffolds. For the first time, they demonstrate that a state-of-the-art model (GPT-5.5) can autonomously accomplish complex optimization tasks without any preset workflow. Across 14 comparisons spanning eight benchmark–target model configurations, OEO achieves 12 wins, 1 tie, and 1 loss, while using only a median of 34.3% of the interaction token budget required by SkillOpt, substantially improving both efficiency and performance.

frontier modelsOpen-Ended Optimizationoptimization process

Outcome-Refining Process Supervision for Code Generation

Dec 19, 2024
ZY
Zhuohao Yu
🏛️ Peking University | Microsoft Research

Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.

Improving code generation for complex programming tasksOvercoming local optima in LLM-generated codeUnifying process and outcome supervision via execution

Latest Papers

What's happening recently
View more

Existing agent evolvers rely on manually designed, fixed search loops and lack autonomous decision-making capabilities. This work proposes a meta-evolution framework that reformulates the optimization process itself as a learnable skill, enabling the evolver to autonomously determine testing, evaluation, and termination strategies. Specifically, the proposed method iteratively refines editable evolution skills by scoring newly generated agents, thereby achieving fully automated and adaptive pipelines. Experimental results demonstrate that this framework yields an average improvement of 13.6 points on primary metrics across multiple benchmarks, significantly outperforming handcrafted approaches. Furthermore, the acquired evolution skills exhibit strong transferability across diverse environments.

Agent EvolutionAutomated Prompt EngineeringMeta-evolution

This study addresses the high computational costs of test-time training for large language models and the credit assignment challenges arising from coupled policy and implementation. We propose Guidance-TTT, a framework that introduces a novel guidance-execution decoupled architecture. Specifically, it freezes a large model to handle code implementation while exclusively training a small model to optimize high-level decision-making policies, thereby restricting test-time learning to short-horizon decision sequences. Furthermore, we design an adaptive group-relative reinforcement learning objective that leverages verifier feedback to enable online policy updates. Experimental results demonstrate that, in offline settings, this framework outperforms state-of-the-art methods across four major domains, including combinatorial optimization and machine learning, while significantly reducing computational overhead and enhancing exploration efficiency.

Computational costCredit assignmentLarge language models

This study addresses the limitation of scalar feedback in automated large language model (LLM) research, where it fails to reveal behavioral conflicts that cause performance stagnation. To overcome this, we propose a competitive behavior feedback mechanism that identifies such conflicts and designs probe metrics to guide code agents in optimizing models. Furthermore, we introduce a reusable ConflictGuide-Skill that integrates literature taxonomy with model evidence to conduct a two-stage evolutionary search for conflict mitigation. Experimental results demonstrate that our approach reduces task error rates by 28% and conflict-related error rates by 14% across five model families, significantly outperforming baseline methods that rely solely on scalar feedback.

AutoResearchcompeting behaviorslarge language models

This study addresses the failure of scientific discovery caused by missing mechanisms under a fixed hypothesis space. We propose an experimental model class revision method that unifies mechanism expressibility and evidence collection into a single sequential decision problem. This approach jointly generates structural edits and diagnostic experiments, triggering theory revision through sequential evidence. To solve this formulation, we introduce a class-level distinguishability objective, anytime-valid sequential evidence, and reinforcement learning-based policies. Evaluated across 400 environments, our method achieves an exact recovery rate of 89.5%, significantly outperforming existing baselines while demonstrating strong transferability. These results establish a new paradigm for automated scientific discovery.

Dynamical SystemsExperimental DesignHypothesis Space Revision

This study addresses the prohibitive costs and inefficiencies associated with model training and validation during AI agent evolution. To this end, we propose an idea-level critic model designed to predict the efficacy of modification proposals and filter high-potential strategies, thereby substituting expensive empirical validation. Methodologically, this specialized model is trained by integrating supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO) reinforcement learning, enabling optimized allocation of validation resources. Experimental results demonstrate that the proposed critic model outperforms existing frontier large language models. Across diverse tasks—including reasoning evolution, continual learning, and policy training—it significantly enhances both evolutionary efficiency under constrained budgets and the quality of final solutions.

AI for MLIdea-level criticsSelf-evolving agents

Hot Scholars

TH

Ting Hua

University of Notre Dame
Efficient learningCompressionReasoning
TD

Tianchen Deng

Shanghai Jiao Tong University
RoboticsComputer Vision
JY

Junwei You

University of Wisconsin-Madison
Autonomous DrivingFoundation ModelsGenerative AIIntelligent Transportation
WX

Weichen Xu

Purdue University
Computer Vision