Score
Design and build pipelines that generate diverse synthetic workflow examples and use them to pretrain models so they learn task-level algorithmic patterns and procedural policies. Implement supervised fine‑tuning and synthetic supervision procedures to produce initial policies or parameterizations that bootstrap generalization across problem instances and serve as starting points for subsequent reinforcement‑learning or task‑specific refinement.
This paper addresses the challenges of structural design and continual optimization for scaffolded language models (LMs) in multi-step tasks. We propose a novel paradigm—*language-supervised training*—that enables non-parametric optimization via natural-language instructions, tool-call trajectories, and human-readable/editable linguistic feedback. This framework dynamically adapts external variables—including prompts, toolchains, and scaffolding code—without modifying model parameters. It is compatible with closed-source LMs, mitigates catastrophic forgetting, and supports human-in-the-loop streaming learning. We introduce the first taxonomy of non-parametric variables tailored to language supervision, unifying prompt engineering, multi-step reasoning orchestration, and feedback modeling. Our approach provides both theoretical foundations and a systematic implementation pathway for deploying hybrid autonomous agents in real-world settings.
While large language models excel at specific tasks, they lack structured, reusable task-level workflows, limiting their reliability, interpretability, and generalization. This work proposes MetaFlow, which formulates workflow generation as a meta-learning problem. Through a two-stage training paradigm—first supervised fine-tuning on synthetically generated data, followed by reinforcement learning with verifiable execution feedback (RLVR)—the model learns to compose operators into general-purpose workflows. MetaFlow achieves, for the first time, zero-shot workflow generation on unseen tasks and novel operator sets using large language models. It attains state-of-the-art performance in a single inference pass across benchmarks in question answering, code generation, and mathematical reasoning, while substantially enhancing cross-task generalization.
Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.
Existing AI research agents often produce seemingly plausible but ineffective machine learning solutions due to a lack of systematic training. To address this, this work proposes the first scalable synthetic task generation framework that automatically constructs high-quality, executable research tasks through topic sampling, proposal generation grounded in real-world Hugging Face datasets, and self-debugging validation. The framework further leverages trajectory distillation—transferring effective research behaviors from GPT-5 to Qwen3—to guide student models in learning valid scientific reasoning paths. Evaluated on the MLGym benchmark, Qwen3-4B and Qwen3-8B models trained with this approach achieve 9% and 12% relative improvements in Area Under the Performance curve (AUP), respectively, substantially outperforming baseline methods.
Supervised fine-tuning (SFT) for subjective open-ended tasks suffers from scarcity of high-quality human annotations and cold-start difficulties in synthetic data generation, as existing workflows rely on annotation-dependent reward models. Method: We propose a reference-free automated synthetic data generation framework built upon an LLM-based evaluator and meta-learning, integrating dynamic task-specific metrics and prompt quality assessment, with Monte Carlo Tree Search (MCTS) enabling self-optimization of the data synthesis pipeline. Contribution/Results: Our key innovation is a hybrid reward mechanism that eliminates dependence on ground-truth annotations. Experiments on educational subjective tasks show models trained on our synthetic data achieve 40–51% performance—substantially outperforming baselines (2–5%)—while reducing manual construction effort from 5–7 hours to just 30 minutes.
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
Web agents exhibit poor adaptability to novel websites, while real-world task data is scarce and existing synthetic data suffers from hallucination, redundancy, and misaligned execution trajectories. Method: We propose a two-stage refinement framework that jointly optimizes synthetic task validity and trajectory fidelity. Our approach integrates webpage-element-classification-guided task generation, runtime task correction, and global-context-aware trajectory refinement to construct high-quality synthetic supervision signals, followed by fine-tuning of open-source web agents. Contribution/Results: Experiments demonstrate substantial improvements in task success rates across multiple unseen websites, outperforming state-of-the-art synthetic data methods. Our framework establishes a scalable, high-fidelity paradigm for synthetic data construction, enhancing web agent generalization under low-resource conditions.
This work investigates how to precisely steer language models toward optimizing behavior on arbitrary differentiable objectives through synthetic training data. The authors propose a reinforcement learning–based approach for synthetic data generation that integrates supervised fine-tuning with higher-order gradient techniques, enabling fine-grained attribution of model outputs for the first time. They introduce a Dataset Policy Gradient (DPG) reward mechanism designed to approximate true gradients without requiring direct access to the model’s internal parameters. This framework allows flexible manipulation of model outputs while remaining agnostic to model internals. The method demonstrates strong empirical performance across diverse and challenging tasks, including embedding QR codes into generated text, producing specified UUIDs, minimizing ℓ² norm of outputs, and performing cross-lingual rewriting.
To address the high performance variance, insufficient success rate (~80%), and fragile sequential execution of policies in contact-intensive robotic assembly tasks under diverse initial conditions, this paper proposes Refinery—a novel framework featuring deployment-time dynamic optimization. Refinery introduces a lightweight online fine-tuning mechanism guided by Bayesian optimization, coupled with Gaussian mixture model-driven sampling of initialization conditions, enabling robust multi-step policy cascading without additional training. The method jointly integrates simulation-to-reality transfer and online adaptation for contact-rich policies, and—crucially—enables runtime adaptive selection of optimal execution conditions for the first time. Experiments demonstrate that Refinery raises the average success rate to 91.51% (+10.98%) in simulation while maintaining comparable performance on physical hardware, successfully completing continuous assembly of up to eight components.
This work addresses the high computational cost of on-policy reinforcement learning for machine learning engineering (MLE) agents, which stems from the need to repeatedly execute full ML pipelines. To overcome this challenge, the authors propose SandMLE, a framework that enables large-scale on-policy reinforcement learning in MLE for the first time. SandMLE leverages multi-agent collaboration to generate synthetic sandbox environments that are structurally complex yet extremely data-efficient—requiring only 50–200 samples per task—thereby preserving real-world task complexity while drastically reducing validation overhead. Experiments demonstrate that SandMLE reduces training time by over 13× and achieves a 20.3%–66.9% improvement in medal rate over supervised fine-tuning baselines on MLE-bench-lite. Furthermore, it attains up to a 32.4% HumanRank generalization gain on MLE-Dojo.
Existing supervised fine-tuning approaches often inherit ineffective steps and logical flaws from teacher trajectories, lacking direct optimization for reasoning validity and trajectory efficiency. This work proposes a bidirectional optimization framework for trajectory generation that, for the first time, converts developer-provided reference patches into implicit process graphs to guide trajectory selection. In the backward phase, it constructs a context-aware factual graph aligned with solution milestones via knowledge distillation; in the forward phase, it scores and prunes teacher trajectories based on this graph, retaining only the shortest valid segments while incorporating a groundedness check to prevent information leakage. Using merely 1.8k curated samples, the method achieves a 10.8 percentage point improvement in Pass@1 on SWE-bench Verified and reduces inference cost by approximately 15%, with consistent gains also observed on SWE-bench Lite.