Synthetic Data Generation for Training Diversified Commonsense Reasoning Models

📅 2026-03-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing commonsense reasoning datasets, which suffer from insufficient scale and diversity, thereby hindering generative models from simultaneously improving response diversity and commonsense plausibility. To overcome this, the authors propose a two-stage synthetic data generation framework that constructs CommonSyn—a large-scale, high-quality dataset tailored for diverse commonsense reasoning—without relying on human annotation, substantially broadening coverage of commonsense scenarios. The approach leverages large language models for both data generation and filtering, and is adaptable to fine-tuning models of varying scales. Experimental results demonstrate that models fine-tuned on CommonSyn significantly outperform baseline systems and those trained on human-annotated data in both response diversity and commonsense quality.

Technology Category

Knowledge Representation and Reasoning: Common-Sense ReasoningNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Conversational agents are required to respond to their users not only with high quality (i.e. commonsense bearing) responses, but also considering multiple plausible alternative scenarios, reflecting the diversity in their responses. Despite the growing need to train diverse commonsense generators, the progress of this line of work has been significantly hindered by the lack of large-scale high-quality diverse commonsense training datasets. Due to the high annotation costs, existing Generative Commonsense Reasoning (GCR) datasets are created using a small number of human annotators, covering only a narrow set of commonsense scenarios. To address this training resource gap, we propose a two-stage method to create the first-ever synthetic dataset CommonSyn for diversified (GCR). The model fine-tuned on our synthetic data jointly increase both generation diversity and quality compared with vanilla models and the model fine-tuned on human-crafted dataset across different size Large Language Models (LLMs)
Problem

Research questions and friction points this paper is trying to address.

synthetic data
commonsense reasoning
response diversity
training dataset
conversational agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic data generation
diversified commonsense reasoning
two-stage method
CommonSyn
large language models
🔎 Similar Papers
No similar papers found.