🤖 AI Summary
This study addresses the cumbersome and fragile nature of dataset construction in zero-data scenarios for AI post-training. To overcome this limitation, it proposes Invent-A-Dataset, a prompt-based system that leverages prompt engineering to directly synthesize large-scale, high-fidelity post-training datasets from textual descriptions under a zero-seed setting. Furthermore, this work systematically evaluates the data synthesis capabilities of five state-of-the-art models. Experimental results demonstrate that the proposed approach improves data quality and diversity by 17% and 19%, respectively, with these advantages becoming significantly more pronounced as the scale increases. Notably, downstream models fine-tuned on the generated data consistently outperform existing baseline methods across all evaluated metrics, establishing the effectiveness of this zero-shot data synthesis paradigm for enhancing post-training performance.
📝 Abstract
Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.