Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the urgent need to systematically evaluate the deceptive capabilities of large language models (LLMs) in high-stakes scenarios characterized by information asymmetry. To this end, the authors introduce ParliamentBench, an open-source evaluation framework based on the social deduction game “Secret Hitler,” which assesses 16 prominent LLMs across 1,600 multi-agent gameplay sessions in tasks involving cooperation, deception, and persuasion. The work presents the first systematic quantification of LLM deception in a controlled strategic environment, proposing three novel metrics to measure social reasoning, logical consistency, and deception stability, alongside releasing a large-scale benchmark dataset of human–AI and AI–AI interactions. Experimental results reveal that while state-of-the-art models (e.g., GPT-5.4, Kimi K2.5) excel at role-playing, most struggle to maintain consistent deception throughout gameplay, with deception consistency rates generally below 50%—and some weaker models performing worse than random (33%) or simple algorithmic (45%) baselines.
📝 Abstract
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.
Problem

Research questions and friction points this paper is trying to address.

deception
large language models
social deduction
reasoning
information asymmetry
Innovation

Methods, ideas, or system contributions that make the work stand out.

deception evaluation
social deduction game
information asymmetry
reasoning consistency
LLM benchmarking