From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding

📅 2025-06-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the scarcity, low diversity, and weak real-world grounding of high-quality instruction data in large language model (LLM) alignment, this paper proposes the *attributed grounding* framework, which synergistically integrates top-down user-context attribution with bottom-up, web-document-driven joint context–instruction generation. Our method employs instruction provenance analysis, multi-granularity context modeling, web retrieval, and structured prompt engineering to construct an end-to-end synthetic pipeline—enabling, for the first time, scalable generation of cognitively inspired yet empirically grounded complex instructions. We release SynthQuestions, a million-scale, high-quality instruction dataset. Evaluated on multiple alignment benchmarks, models trained on SynthQuestions achieve significant performance gains, with improvements consistently scaling with the volume of underlying web corpora.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
The pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models (LLMs). While there are methods capable of generating synthetic instructions at scale, they either suffer from limited grounding sources, leading to a narrow distribution, or rely on trivial extensions that fail to produce meaningful trajectories in terms of complexity. In contrast, instructions that benefit efficient alignment are typically crafted with cognitive insights and grounded in real-world use cases. In this paper, we synthesize such instructions using attributed grounding, which involves 1) a top-down attribution process that grounds a selective set of real instructions to situated users, and 2) a bottom-up synthesis process that leverages web documents to first generate a situation, then a meaningful instruction. This framework allows us to harvest diverse and complex instructions at scale, utilizing the vast range of web documents. Specifically, we construct a dataset of 1 million instructions, called SynthQuestions, and demonstrate that models trained on it achieve leading performance on several common benchmarks, with improvements that continually scale with more web corpora. Data, models and codes will be available at https://github.com/Ignoramus0817/SynthQuestions.
Problem

Research questions and friction points this paper is trying to address.

Generating diverse synthetic instructions for LLM alignment
Overcoming limited grounding sources in instruction synthesis
Enhancing instruction complexity with attributed grounding methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Top-down attribution grounds real instructions
Bottom-up synthesis leverages web documents
Generates diverse complex instructions at scale
🔎 Similar Papers
No similar papers found.