synthetic workflow pretraining

Design and build pipelines that generate diverse synthetic workflow examples and use them to pretrain models so they learn task-level algorithmic patterns and procedural policies. Implement supervised fine‑tuning and synthetic supervision procedures to produce initial policies or parameterizations that bootstrap generalization across problem instances and serve as starting points for subsequent reinforcement‑learning or task‑specific refinement.

syntheticworkflowpretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

While large language models excel at specific tasks, they lack structured, reusable task-level workflows, limiting their reliability, interpretability, and generalization. This work proposes MetaFlow, which formulates workflow generation as a meta-learning problem. Through a two-stage training paradigm—first supervised fine-tuning on synthetically generated data, followed by reinforcement learning with verifiable execution feedback (RLVR)—the model learns to compose operators into general-purpose workflows. MetaFlow achieves, for the first time, zero-shot workflow generation on unseen tasks and novel operator sets using large language models. It attains state-of-the-art performance in a single inference pass across benchmarks in question answering, code generation, and mathematical reasoning, while substantially enhancing cross-task generalization.

large language modelsstructural consistencytask-level patterns

Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.

execution traceLLM agentssupervision data

Existing AI research agents often produce seemingly plausible but ineffective machine learning solutions due to a lack of systematic training. To address this, this work proposes the first scalable synthetic task generation framework that automatically constructs high-quality, executable research tasks through topic sampling, proposal generation grounded in real-world Hugging Face datasets, and self-debugging validation. The framework further leverages trajectory distillation—transferring effective research behaviors from GPT-5 to Qwen3—to guide student models in learning valid scientific reasoning paths. Evaluated on the MLGym benchmark, Qwen3-4B and Qwen3-8B models trained with this approach achieve 9% and 12% relative improvements in Area Under the Performance curve (AUP), respectively, substantially outperforming baseline methods.

Agent TrainingAI ScientistAutomatic Scientific Discovery

AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search

Nov 12, 2025
SB
Shuzhen Bi
🏛️ Shanghai Innavation Institute | Shanghai Institute of AI for Education | East China Normal University

Supervised fine-tuning (SFT) for subjective open-ended tasks suffers from scarcity of high-quality human annotations and cold-start difficulties in synthetic data generation, as existing workflows rely on annotation-dependent reward models. Method: We propose a reference-free automated synthetic data generation framework built upon an LLM-based evaluator and meta-learning, integrating dynamic task-specific metrics and prompt quality assessment, with Monte Carlo Tree Search (MCTS) enabling self-optimization of the data synthesis pipeline. Contribution/Results: Our key innovation is a hybrid reward mechanism that eliminates dependence on ground-truth annotations. Experiments on educational subjective tasks show models trained on our synthetic data achieve 40–51% performance—substantially outperforming baselines (2–5%)—while reducing manual construction effort from 5–7 hours to just 30 minutes.

Automating synthetic dataset generation for LLM fine-tuning without reference datasetsReducing human effort in creating high-quality training data for specialized tasksSolving cold start problem in workflow optimization for subjective tasks

Outcome-Refining Process Supervision for Code Generation

Dec 19, 2024
ZY
Zhuohao Yu
🏛️ Peking University | Microsoft Research

Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.

Improving code generation for complex programming tasksOvercoming local optima in LLM-generated codeUnifying process and outcome supervision via execution

Latest Papers

What's happening recently
View more

Adapting Web Agents with Synthetic Supervision

Nov 08, 2025
ZW
Zhaoyang Wang
🏛️ UNC-Chapel Hill | Purdue University | Microsoft

Web agents exhibit poor adaptability to novel websites, while real-world task data is scarce and existing synthetic data suffers from hallucination, redundancy, and misaligned execution trajectories. Method: We propose a two-stage refinement framework that jointly optimizes synthetic task validity and trajectory fidelity. Our approach integrates webpage-element-classification-guided task generation, runtime task correction, and global-context-aware trajectory refinement to construct high-quality synthetic supervision signals, followed by fine-tuning of open-source web agents. Contribution/Results: Experiments demonstrate substantial improvements in task success rates across multiple unseen websites, outperforming state-of-the-art synthetic data methods. Our framework establishes a scalable, high-fidelity paradigm for synthetic data construction, enhancing web agent generalization under low-resource conditions.

SynthAgent refines synthetic tasks and trajectories to improve data qualitySynthetic data generation suffers from hallucinations and noisy trajectoriesWeb agents struggle to adapt to new websites due to data scarcity

This work investigates how to precisely steer language models toward optimizing behavior on arbitrary differentiable objectives through synthetic training data. The authors propose a reinforcement learning–based approach for synthetic data generation that integrates supervised fine-tuning with higher-order gradient techniques, enabling fine-grained attribution of model outputs for the first time. They introduce a Dataset Policy Gradient (DPG) reward mechanism designed to approximate true gradients without requiring direct access to the model’s internal parameters. This framework allows flexible manipulation of model outputs while remaining agnostic to model internals. The method demonstrates strong empirical performance across diverse and challenging tasks, including embedding QR codes into generated text, producing specified UUIDs, minimizing ℓ² norm of outputs, and performing cross-lingual rewriting.

differentiable metriclanguage modelsmodel control

Refinery: Active Fine-tuning and Deployment-time Optimization for Contact-Rich Policies

Oct 13, 2025
BT
Bingjie Tang
🏛️ University of Southern California | NVIDIA Corporation | University of Washington | University of Sydney

To address the high performance variance, insufficient success rate (~80%), and fragile sequential execution of policies in contact-intensive robotic assembly tasks under diverse initial conditions, this paper proposes Refinery—a novel framework featuring deployment-time dynamic optimization. Refinery introduces a lightweight online fine-tuning mechanism guided by Bayesian optimization, coupled with Gaussian mixture model-driven sampling of initialization conditions, enabling robust multi-step policy cascading without additional training. The method jointly integrates simulation-to-reality transfer and online adaptation for contact-rich policies, and—crucially—enables runtime adaptive selection of optimal execution conditions for the first time. Experiments demonstrate that Refinery raises the average success rate to 91.51% (+10.98%) in simulation while maintaining comparable performance on physical hardware, successfully completing continuous assembly of up to eight components.

Enabling reliable policy chaining for multi-part assembliesImproving low success rates of robotic assembly policiesReducing performance variance across diverse initial conditions

This work addresses the high computational cost of on-policy reinforcement learning for machine learning engineering (MLE) agents, which stems from the need to repeatedly execute full ML pipelines. To overcome this challenge, the authors propose SandMLE, a framework that enables large-scale on-policy reinforcement learning in MLE for the first time. SandMLE leverages multi-agent collaboration to generate synthetic sandbox environments that are structurally complex yet extremely data-efficient—requiring only 50–200 samples per task—thereby preserving real-world task complexity while drastically reducing validation overhead. Experiments demonstrate that SandMLE reduces training time by over 13× and achieves a 20.3%–66.9% improvement in medal rate over supervised fine-tuning baselines on MLE-bench-lite. Furthermore, it attains up to a 32.4% HumanRank generalization gain on MLE-Dojo.

agent verificationcomputational bottleneckmachine learning engineering

Existing supervised fine-tuning approaches often inherit ineffective steps and logical flaws from teacher trajectories, lacking direct optimization for reasoning validity and trajectory efficiency. This work proposes a bidirectional optimization framework for trajectory generation that, for the first time, converts developer-provided reference patches into implicit process graphs to guide trajectory selection. In the backward phase, it constructs a context-aware factual graph aligned with solution milestones via knowledge distillation; in the forward phase, it scores and prunes teacher trajectories based on this graph, retaining only the shortest valid segments while incorporating a groundedness check to prevent information leakage. Using merely 1.8k curated samples, the method achieves a 10.8 percentage point improvement in Pass@1 on SWE-bench Verified and reduces inference cost by approximately 15%, with consistent gains also observed on SWE-bench Lite.

reference patchsoftware-engineering agentssupervised fine-tuning

Hot Scholars

JM

Jelmer M. Wolterink

Associate Professor, University of Twente
Deep LearningMedical Image AnalysisArtificial IntelligenceMathematics
RZ

Rui Zhong

Hokkaido University, Information Initiative Center, Specially Appointed Assistant Professor
Computational IntelligenceEvolutionary ComputationLarge-scale OptimizationMeta/Hyper-heuristics
CX

Caiming Xiong

Salesforce Research
Machine LearningNLPComputer VisionMultimedia
HC

Hua Chen

Assistant Professor, ZJU-UIUC Institute; Co-founder, LimX Dynamics
RoboticsEmbodied AIRobot LearningReinforcement Learning
HK

Hyunjae Kim

Yale University
Natural Language ProcessingBiomedical InformaticsHealthcare