It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited reasoning capabilities of pretraining data and the unclear mechanisms underlying synthetic data. To this end, it introduces SYNTH, an open-source synthetic corpus generated through structured data augmentation seeded from Wikipedia. The proposed approach pioneers a fully synthetic, single-stage training paradigm that unifies pretraining, mid-training, and post-training, thereby significantly reducing reliance on large-scale web-crawled data. Experiments demonstrate that, under equivalent computational budgets, this method outperforms traditional web data across dense and Mixture-of-Experts (MoE) models ranging from 56M to 13B parameters, achieving high factual accuracy and enhanced performance for smaller models with minimal token consumption. Both the SYNTH dataset and the Baguettotron model family have been publicly released.
📝 Abstract
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.
Problem

Research questions and friction points this paper is trying to address.

synthetic data
pre-training datasets
large language models
data efficiency
knowledge acquisition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data
Single-Stage Training
Data Efficiency
Large Language Models
Open-source Corpus
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pierre-Carl Langlais
1PleIAs; 2Sorbonne Center for Artificial Intelligence; 3Sciences Po Médialab
Pieter Delobelle
Pieter Delobelle
KU Leuven
machine learningNLPfairnessAI ethics
Y
Yannick Detrois
1PleIAs; 4EPFL
Pavel Chizhov
Pavel Chizhov
Researcher and Ph.D. Student at cairo.thws
natural language processinglarge language modelscomputer visionself-supervised learning
C
Carlos Rosas-Hinostroza
1PleIAs; 7Lattice, ENS-PSL
N
Neil Si Smail
1PleIAs
B
Benjamin Burtin
1PleIAs
H
Hanna Shcharbakova
8TU Munich, Munich Center for Machine Learning
I
Ivan Yamshchikov
1PleIAs; 5CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt
A
Anastasia Stasenko
1PleIAs; 6Paris Dauphine-PSL