automated data construction

Designs and implements automated pipelines and tooling that programmatically assemble, collect, label, augment, and curate training datasets, including scraping or instrumentation, metadata capture, and workflow orchestration. Builds synthetic-data generators and composition systems that produce paired reference–target composites and diverse insertion examples, and instruments quality/diversity checks to monitor and improve the produced training data.

automateddataconstruction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.

execution traceLLM agentssupervision data

Meta-Learning and Synthetic Data for Automated Pretraining and Finetuning

Jun 11, 2025
FF
Fabio Ferreira
🏛️ Albert-Ludwigs-Universität Freiburg

In the era of large language models, selecting and fine-tuning pre-trained models remains challenging due to the absence of efficient adaptation mechanisms across vast model zoos and severe limitations in labeled data, hindering few-shot generalization. Method: (1) We systematically integrate meta-learning into the deep learning pipeline, constructing a task-prior-driven pipeline ranking surrogate model; (2) we quantitatively characterize the critical role of data augmentation in self-supervised learning; (3) we propose a differentiable neural synthetic data generator, replacing conventional reinforcement learning–based approaches. Contribution/Results: Our framework significantly outperforms human-crafted fine-tuning baselines on standard CV and NLP benchmarks. It achieves substantial gains in few-shot settings and enables zero-shot cross-environment generalization of the synthetic data generator, demonstrating robust adaptability without domain-specific retraining.

Automating deep learning pipeline selection for Computer Vision tasksMeta-learning data augmentation to improve Self-Supervised LearningUsing synthetic data to enhance Reinforcement Learning environments

Current data processing pipelines for post-training large language models—encompassing cleaning, deduplication, synthesis, and quality filtering—are fragmented and lack auditability and sample-level decision transparency. This work proposes the first end-to-end configurable data processing framework that unifies data ingestion, cleaning, LLM-driven synthesis across eight task types, three-tiered quality gating, and export modules. The system introduces sample-level provenance tracking and a precise hallucination verification mechanism. It supports six input formats and over 100 model APIs via LiteLLM, offering both a YAML-driven command-line interface and a Python API. Outputs are compatible with five training formats used by TRL, Unsloth, and AlignTune, substantially enhancing transparency, reproducibility, and scalability in post-training data preparation.

data curationLLM post-trainingpipeline auditability

ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs

Dec 17, 2025
HC
Hao Chen
🏛️ North China University of Technology | Meituan | Institute of Software, Chinese Academy of Sciences | National University of Singapore | East China Normal University

Existing tool learning approaches rely on real API calls, incurring high computational costs, exhibiting poor generalization, and lacking multi-hop reasoning and self-reflection capabilities. Method: We propose the first real-API-call-free framework for synthesizing multi-hop search tool learning data. Given (question, gold context, answer) triplets, it automatically generates high-quality, diverse training data via lightweight virtual tool modeling. Our method innovatively integrates multi-hop reasoning chain generation with self-reflection enhancement, and establishes a multi-layer verification system—combining rule-based and model-based checks—to ensure data fidelity. Contribution/Results: Experiments demonstrate that an 8B-parameter model trained on our synthetic data surpasses GPT-4o across multiple benchmarks. To foster reproducibility and community advancement, we publicly release both the code and the dataset.

Enables multi-hop search with self-reflection mechanismsSynthesizes tool-learning data without real API callsTrains smaller models to outperform larger ones on benchmarks

CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation

Sep 03, 2024
IZ
Ingo Ziegler
🏛️ University of Copenhagen | Center for Information and Language Processing (CIS) | LMU Munich

Addressing the challenge of constructing high-quality, domain-specific annotated data—often costly and labor-intensive—this paper proposes a few-shot-driven synthetic data generation paradigm. Given only a small set of user-provided examples, the method retrieves semantically relevant real-world text from large-scale web corpora and leverages instruction-tuned large language models (LLMs) to automatically generate well-formatted, task-specific synthetic training data. It is the first approach to synergistically integrate corpus retrieval with LLM-based augmentation, enabling zero human annotation, domain adaptability, and efficient few-shot generalization. Empirical evaluation across biomedical, medical, and commonsense question answering (QA), as well as summarization tasks, demonstrates that models trained on the generated data achieve a 46-point preference score improvement over human-annotated baselines in summarization, while QA models match or surpass the performance of general-purpose foundation models.

Generates synthetic datasets for specialized tasks efficientlyOutperforms human-curated and other synthetic data methodsUses corpus retrieval and LLM augmentation for customization

Latest Papers

What's happening recently
View more

Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.

benchmarkingdata qualityevaluation framework

Existing synthetic data generation tools often suffer from complex workflows, inconsistent standards, and limited cross-modal extensibility, hindering their ability to meet the data demands of large language models in specialized domains and low-resource languages. This work proposes a configuration-driven, end-to-end open-source framework that standardizes multi-source data synthesis through a unified and controllable paradigm. Featuring a highly modular architecture, the framework flexibly adapts to diverse tasks and supports high-quality data generation across multiple pathways, modalities, and languages. Integrated with both a graphical user interface and command-line utilities, it significantly lowers the barrier to entry for users. Empirical evaluations demonstrate that the framework effectively balances generation efficiency and data quality across various scenarios, thereby accelerating the practical deployment of synthetic data in model training pipelines.

data scarcitymultilingualmultimodal

This work addresses the heavy reliance on expert knowledge in designing and debugging scientific workflows, a challenge exacerbated by existing large language model approaches that directly generate code without ensuring transparency, reproducibility, or seamless system integration. To overcome these limitations, we propose an AI-assisted scientific workflow management framework that decouples user intent from implementation through a structured specification phase, enabling specification-driven workflow generation and validation. We further introduce a multi-layer debugging agent powered by large language models to automate error diagnosis and correction. By deeply integrating with the Pegasus workflow system via the Model Context Protocol (MCP), our approach supports end-to-end workflow lifecycle management. Empirical evaluation demonstrates successful generation and execution of federated learning medical imaging workflows comprising thousands of tasks, substantially reducing debugging effort and empowering non-expert users to construct complex workflows adhering to expert-level design patterns.

debugginglarge language modelsreproducibility

Existing coding agents are largely confined to code generation and lack support for the full workflow lifecycle, including composition, iteration, deployment, and sharing. This work proposes CURATE, a novel system that integrates modular cataloging and FAIR principles into a large language model–driven multi-agent framework to enable human-in-the-loop, end-to-end workflow development and automated execution. Built upon Claude Opus 4.8, CURATE incorporates user-in-the-loop mechanisms and a module registry to facilitate cross-workflow sharing of reusable components. The system successfully reproduces four SeBS-Flow benchmark workflows and automatically constructs a complex anaerobic digestion simulation pipeline, demonstrating its feasibility and effectiveness in supporting comprehensive, collaborative scientific workflow automation.

code generationdeploymentmodule reuse

This work addresses the lack of quantitative evaluation in existing methods regarding how generated data affects downstream model performance, which hinders reliable synthetic data quality assurance. The authors propose a model-aware synthetic data generation framework that, for the first time, leverages acquisition functions from active learning as interpretable, model-centric reward signals. Integrating reinforcement learning–based generation, a rejection-sampling alternative strategy, and generalization techniques across models and resource scales, the framework guides language models to produce data with higher information content and greater task impact. Experiments on mathematical reasoning, medical question answering, and code generation demonstrate that student models trained on the generated data achieve performance gains of 2–7% and exhibit significantly improved robustness against catastrophic forgetting.

acquisition functionsdata qualitydownstream learner impact

Hot Scholars

QM

Qiang Ma

Assistant Researcher of Tsinghua University
wireless sensor networksnetwork diagnosis
PL

Ping Luo

National University of Defense Technology
distributed_computing
TE

Tome Eftimov

Computer Systems Department, Jožef Stefan Institute
StatisticsStochastic Optimization AlgorithmsMachine learningNatural Language Processing
JL

Jian Luan

Toshiba, Microsoft, Xiaomi
LLMVLMTTSSinging Synthesis
QZ

Qiaosheng Zhang

Department of Anesthesiology, New York University School of Medcine
NeuroscienceNeural EngineeringPain CircuitryBrain Machine Interface