🤖 AI Summary
This study evaluates the reasoning capabilities of large language models (LLMs) in democratic deliberation scenarios that lack objective answers and require integration of diverse value systems, thereby exposing the limitations of current evaluation frameworks grounded in verifiable tasks and procedural metrics. For the first time, the empirically validated Deliberative Reasoning Index (DRI) from political science is adapted to assess multi-agent LLM deliberations. The authors conduct a systematic analysis of 1,980 dialogues across 12 civic issues involving 11 state-of-the-art models. Results reveal that while LLM groups achieve procedural discourse quality comparable to humans, they exhibit less than one-third the level of perspective diversity. Moreover, on ethically contentious topics, LLMs struggle to reach consensus, with deliberation often exacerbating disagreement. Role-based prompting fails to replicate human-like deliberative dynamics, challenging the assumption that LLMs can function as autonomous deliberative agents.
📝 Abstract
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.