Score
Design and build model pretraining pipelines and synthetic-data generation processes that create simulated priors and diverse synthetic sequences, and pretrain models on these datasets so they learn broad structural patterns such as posterior-predictive behavior and item-to-item transitions. Evaluate and engineer methods for transferring these pretrained models to real data, including synthetic-to-real fine-tuning, prior-data pretraining, and approaches that enable cross-domain or parameter-free inference.
In the era of large language models, selecting and fine-tuning pre-trained models remains challenging due to the absence of efficient adaptation mechanisms across vast model zoos and severe limitations in labeled data, hindering few-shot generalization. Method: (1) We systematically integrate meta-learning into the deep learning pipeline, constructing a task-prior-driven pipeline ranking surrogate model; (2) we quantitatively characterize the critical role of data augmentation in self-supervised learning; (3) we propose a differentiable neural synthetic data generator, replacing conventional reinforcement learning–based approaches. Contribution/Results: Our framework significantly outperforms human-crafted fine-tuning baselines on standard CV and NLP benchmarks. It achieves substantial gains in few-shot settings and enables zero-shot cross-environment generalization of the synthetic data generator, demonstrating robust adaptability without domain-specific retraining.
研究通过比较合成数据生成器与基准数据集的结构描述符,探讨了合成预训练先验对表格基础模型下游任务的支持程度。
Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.
This study investigates the fundamental limitations of synthetic data augmentation in enhancing sample information for statistical inference, with particular emphasis on its theoretical constraints when incorporating prior knowledge. Treating synthetic data as a model of the prior, we formally define the synthetic distribution within both maximum likelihood and Bayesian frameworks. By integrating Fisher information, information theory, and statistical decision theory, we establish an intrinsic upper bound on the marginal information gain achievable through synthetic data. Our analysis reveals that naive prior specifications lack epistemic justification and generally fail to improve conventional inferential performance. Nevertheless, under the training/test partition paradigm, synthetic data can effectively regularize high-dimensional model spaces by imposing structural constraints, thereby serving a beneficial regularization role despite its limited informational contribution.
This study addresses the unclear relationship between properties of synthetic data and model learning of relational structures, which currently disconnects performance improvements from their underlying mechanisms. To bridge this gap, we construct a causal chain spanning data properties, learning mechanisms, and downstream behaviors. By integrating Relational Transformers, data attribution analysis, and foreign key intervention experiments, we systematically trace the mapping pathway from pretraining data to model behavior. Our findings reveal that the necessity for cross-table prediction is the core factor driving relational computation. Furthermore, we demonstrate that the RelDiff generator establishes its advantage by inducing unique serial cross-table pathways; disrupting this mechanism entirely eliminates its performance gains. These insights provide explicit theoretical guidance for the principled design of synthetic relational data.
The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.
Existing time series pretraining methods struggle to generalize effectively across multiple datasets due to discrepancies in input length and channel dimensions. This work proposes ADAPT, a novel pretraining paradigm that enables unified modeling across 162 time series classification datasets by adaptively aligning the physical attributes of time series data. Integrating self-supervised learning with a hybrid batch training strategy, ADAPT overcomes the generalization limitations inherent in conventional many-to-one pretraining approaches. The method achieves state-of-the-art performance on multiple benchmarks, establishing a foundational framework for developing general-purpose foundation models for time series analysis.
This study addresses the limited reasoning capabilities of pretraining data and the unclear mechanisms underlying synthetic data. To this end, it introduces SYNTH, an open-source synthetic corpus generated through structured data augmentation seeded from Wikipedia. The proposed approach pioneers a fully synthetic, single-stage training paradigm that unifies pretraining, mid-training, and post-training, thereby significantly reducing reliance on large-scale web-crawled data. Experiments demonstrate that, under equivalent computational budgets, this method outperforms traditional web data across dense and Mixture-of-Experts (MoE) models ranging from 56M to 13B parameters, achieving high factual accuracy and enhanced performance for smaller models with minimal token consumption. Both the SYNTH dataset and the Baguettotron model family have been publicly released.
This work addresses the risks of bias in statistical inference when using synthetic data generated by modern generative AI models—such as diffusion models, GANs, and large language models—due to model misspecification, underestimation of uncertainty, and insufficient generalization. It presents the first systematic integration of generative modeling with statistical inference theory, clarifying the assumptions and conditions under which synthetic data can reliably support scientific discovery. By unifying uncertainty quantification, model diagnostics, and downstream task analysis, the study proposes a principled framework that delineates effective usage guidelines, identifies critical failure modes, and offers practical recommendations for researchers and developers, along with directions for future research.
This study systematically investigates the effectiveness of synthetic data in time series forecasting and its dependence on model architecture. Drawing on 4,218 experiments across nine configurations, the authors evaluate five prominent deep learning models—including TimesNet, iTransformer, DLinear, and PatchTST—on four synthetic signals and seven real-world datasets. The work reveals, for the first time, that the benefits of synthetic data are highly architecture-dependent: channel-mixing models consistently gain substantial improvements, particularly under low-data regimes, whereas synthetic data proves detrimental in 67% of all experimental settings. The study proposes effective usage strategies tailored to channel-mixing architectures and demonstrates that progressive scheduling outperforms hard curriculum switching. Among synthetic generation methods, only the seasonal-trend decomposition generator yields consistent performance gains.