AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search

📅 2025-11-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Supervised fine-tuning (SFT) for subjective open-ended tasks suffers from scarcity of high-quality human annotations and cold-start difficulties in synthetic data generation, as existing workflows rely on annotation-dependent reward models. Method: We propose a reference-free automated synthetic data generation framework built upon an LLM-based evaluator and meta-learning, integrating dynamic task-specific metrics and prompt quality assessment, with Monte Carlo Tree Search (MCTS) enabling self-optimization of the data synthesis pipeline. Contribution/Results: Our key innovation is a hybrid reward mechanism that eliminates dependence on ground-truth annotations. Experiments on educational subjective tasks show models trained on our synthetic data achieve 40–51% performance—substantially outperforming baselines (2–5%)—while reducing manual construction effort from 5–7 hours to just 30 minutes.

Technology Category

Natural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)Search and Optimization: Metareasoning and Metaheuristics

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Supervised fine-tuning (SFT) of large language models (LLMs) for specialized tasks requires high-quality datasets, but manual curation is prohibitively expensive. Synthetic data generation offers scalability, but its effectiveness relies on complex, multi-stage workflows, integrating prompt engineering and model orchestration. Existing automated workflow methods face a cold start problem: they require labeled datasets for reward modeling, which is especially problematic for subjective, open-ended tasks with no objective ground truth. We introduce AutoSynth, a framework that automates workflow discovery and optimization without reference datasets by reframing the problem as a Monte Carlo Tree Search guided by a novel dataset-free hybrid reward. This reward enables meta-learning through two LLM-as-judge components: one evaluates sample quality using dynamically generated task-specific metrics, and another assesses workflow code and prompt quality. Experiments on subjective educational tasks show that while expert-designed workflows achieve higher human preference rates (96-99% win rates vs. AutoSynth's 40-51%), models trained on AutoSynth-generated data dramatically outperform baselines (40-51% vs. 2-5%) and match or surpass expert workflows on certain metrics, suggesting discovery of quality dimensions beyond human intuition. These results are achieved while reducing human effort from 5-7 hours to just 30 minutes (>90% reduction). AutoSynth tackles the cold start issue in data-centric AI, offering a scalable, cost-effective method for subjective LLM tasks. Code: https://github.com/bisz9918-maker/AutoSynth.
Problem

Research questions and friction points this paper is trying to address.

Automating synthetic dataset generation for LLM fine-tuning without reference datasets
Solving cold start problem in workflow optimization for subjective tasks
Reducing human effort in creating high-quality training data for specialized tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automates workflow discovery via Monte Carlo Tree Search
Uses dataset-free hybrid reward with LLM-as-judge components
Reduces human effort by over 90% for subjective tasks
S
Shuzhen Bi
Shanghai Innavation Institute, Shanghai, 200231, China
Chang Song
Chang Song
Shanghai Institute of AI for Education, East China Normal University, Shanghai, 200062, China
S
Siyu Song
Shanghai Institute of AI for Education, East China Normal University, Shanghai, 200062, China
J
Jinze Lv
School of Computer Science and Technology, East China Normal University, Shanghai, 200062, China
J
Jian Chen
Shanghai Institute of AI for Education, East China Normal University, Shanghai, 200062, China
X
Xinyun Wang
Department of Education Information Technology, East China Normal University, Shanghai, 200062, China
A
Aimin Zhou
Shanghai Innavation Institute, Shanghai, 200231, China; Shanghai Institute of AI for Education, East China Normal University, Shanghai, 200062, China; School of Computer Science and Technology, East China Normal University, Shanghai, 200062, China
H
Hao Hao
Shanghai Institute of AI for Education, East China Normal University, Shanghai, 200062, China