D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

πŸ“… 2026-07-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current benchmarks for evaluating value alignment in large language models exhibit insufficient coverage and oversimplified formulations when addressing everyday scenarios involving multiple, often conflicting values. To address this gap, this work proposes D2VBench, a comprehensive evaluation benchmark comprising 10,000 real-world moral dilemma instances grounded in 158 fine-grained value concepts. Constructed through a human-AI collaborative, multi-stage pipeline, D2VBench employs a hybrid assessment paradigm that integrates both multiple-choice and open-ended questions. It is the first benchmark to systematically capture multidimensional value conflicts encountered in daily life. Empirical evaluations across eight mainstream large language models demonstrate that D2VBench achieves high reliability and robustness, effectively characterizing models’ alignment capabilities across diverse value dimensions.
πŸ“ Abstract
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment. To address these issues, we propose D2VBench, a value alignment benchmark comprising 10,000 instances of real daily dilemma scenarios constructed through a multi-stage collaboration between LLMs and humans, grounded in 158 manually annotated fine-grained value concepts. For evaluation on the benchmark, we present a hybrid evaluation paradigm that integrates multiple-choice questions with open-ended questions. We conduct comprehensive evaluations on eight mainstream LLMs. Experimental results demonstrate that D2VBench exhibits high reliability and robustness, effectively reflecting the LLMs' alignment across different value categories and dimensions, and providing a more realistic and fine-grained tool for research on value alignment. The dataset is available at https://github.com/tjunlp-lab/D2VBench.
Problem

Research questions and friction points this paper is trying to address.

value alignment
large language models
value dilemmas
evaluation benchmark
daily scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

value alignment
daily dilemma scenarios
fine-grained values
hybrid evaluation paradigm
LLM benchmarking
S
Siyi Hao
TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China
Y
Yidi Cao
The International Joint Institute of Tianjin University, Fuzhou, China
L
Linhao Yu
TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China
Y
Yuqi Ren
TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China
Deyi Xiong
Deyi Xiong
Professor, College of Intelligence and Computing, Tianjin University, China
Natural Language ProcessingLarge Language ModelsAI4Science