🤖 AI Summary
To address the scarcity, low diversity, and weak real-world grounding of high-quality instruction data in large language model (LLM) alignment, this paper proposes the *attributed grounding* framework, which synergistically integrates top-down user-context attribution with bottom-up, web-document-driven joint context–instruction generation. Our method employs instruction provenance analysis, multi-granularity context modeling, web retrieval, and structured prompt engineering to construct an end-to-end synthetic pipeline—enabling, for the first time, scalable generation of cognitively inspired yet empirically grounded complex instructions. We release SynthQuestions, a million-scale, high-quality instruction dataset. Evaluated on multiple alignment benchmarks, models trained on SynthQuestions achieve significant performance gains, with improvements consistently scaling with the volume of underlying web corpora.
📝 Abstract
The pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models (LLMs). While there are methods capable of generating synthetic instructions at scale, they either suffer from limited grounding sources, leading to a narrow distribution, or rely on trivial extensions that fail to produce meaningful trajectories in terms of complexity. In contrast, instructions that benefit efficient alignment are typically crafted with cognitive insights and grounded in real-world use cases. In this paper, we synthesize such instructions using attributed grounding, which involves 1) a top-down attribution process that grounds a selective set of real instructions to situated users, and 2) a bottom-up synthesis process that leverages web documents to first generate a situation, then a meaningful instruction. This framework allows us to harvest diverse and complex instructions at scale, utilizing the vast range of web documents. Specifically, we construct a dataset of 1 million instructions, called SynthQuestions, and demonstrate that models trained on it achieve leading performance on several common benchmarks, with improvements that continually scale with more web corpora. Data, models and codes will be available at https://github.com/Ignoramus0817/SynthQuestions.