A Unified Moral-Value Dataset for Instruction Tuning

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the insufficient alignment of current large language models with human moral values and the scarcity of instruction-tuning data specifically covering ethical scenarios. To bridge this gap, the authors present the first unified, comprehensive moral-values dataset by harmonizing multiple heterogeneous ethical data sources into a standardized instruction-response format. They systematically investigate strategies for interleaving this moral dataset with general-purpose instruction data during fine-tuning. Experimental results demonstrate that, under an appropriate mixing ratio, this approach significantly enhances model performance on value-oriented tasks without compromising general capabilities, thereby revealing the critical role of data mixture proportions in achieving effective moral alignment.
📝 Abstract
Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human values is still an open problem. Recent studies show that instruction tuning has strong potential for zero-shot tasks and may serve as an effective approach to addressing value alignment. Nevertheless, although many datasets for instruction tuning already exist, they are not specifically designed around moral scenarios and behaviors. We construct a unified moral-value dataset that can be directly used for instruction tuning. This dataset is built upon existing moral-value datasets by merging them into a unified corpus and converting them into an instruction-response format. We show that training on a mixed dataset combining general task datasets with our dataset preserves general-task performance, and we report preliminary observations on how the mixing ratio affects value-oriented task performance. Our work provides a moral-value dataset for instruction tuning and offers a useful resource for further alignment research. The dataset is available at https://huggingface.co/datasets/teohzzh/value-for-instruction-tuning.
Problem

Research questions and friction points this paper is trying to address.

value alignment
instruction tuning
moral values
large language models
dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

moral-value alignment
instruction tuning
unified dataset
value-oriented tasks
large language models