A dataset of rated conceptual arguments

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating large language models’ reasoning capabilities on conceptual questions—such as those in philosophy, AI safety, and decision theory—that lack objective ground truth. The authors present the first systematically constructed and annotated dataset comprising 951 argumentative commentaries derived from 442 stance-based texts, spanning domains including AI safety, ethics, and political theory. They introduce a multidimensional expert scoring framework assessing centrality, strength, correctness, and clarity, yielding 1,458 annotations, and propose two scoring functions that decompose holistic conclusions into evaluable local argument units. Benchmarking experiments demonstrate a strong correlation between model performance and general language capabilities, thereby validating the dataset’s effectiveness and practicality as a novel benchmark for conceptual reasoning evaluation.
📝 Abstract
Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exists, but on which progress can still be made by debating arguments. Most philosophical questions are of this kind, as are central components of questions in AI safety, decision theory, and social choice. Our approach is based on the view that while bottom-line conclusions on such questions are hard to evaluate, individual contextualized arguments can be evaluated far more reliably. We therefore introduce a dataset of 951 argumentative critiques of 442 position texts, spanning topics from AI safety and decision theory to ethics and politics, with 1,458 ratings by six expert raters along dimensions including centrality, strength, correctness, and clarity. We propose two scoring functions and benchmark a range of models. Performance tracks general capability rankings.
Problem

Research questions and friction points this paper is trying to address.

conceptual questions
argument evaluation
AI safety
decision theory
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

conceptual reasoning
argument evaluation
expert-rated dataset
AI safety
debate-based assessment
🔎 Similar Papers
No similar papers found.