🤖 AI Summary
This study addresses the urgent need to systematically evaluate the deceptive capabilities of large language models (LLMs) in high-stakes scenarios characterized by information asymmetry. To this end, the authors introduce ParliamentBench, an open-source evaluation framework based on the social deduction game “Secret Hitler,” which assesses 16 prominent LLMs across 1,600 multi-agent gameplay sessions in tasks involving cooperation, deception, and persuasion. The work presents the first systematic quantification of LLM deception in a controlled strategic environment, proposing three novel metrics to measure social reasoning, logical consistency, and deception stability, alongside releasing a large-scale benchmark dataset of human–AI and AI–AI interactions. Experimental results reveal that while state-of-the-art models (e.g., GPT-5.4, Kimi K2.5) excel at role-playing, most struggle to maintain consistent deception throughout gameplay, with deception consistency rates generally below 50%—and some weaker models performing worse than random (33%) or simple algorithmic (45%) baselines.
📝 Abstract
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.