Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing lie detection probes in large language models to spurious correlations such as instruction following, which impedes the identification of false statements contradicting real-world beliefs during role-playing. We present the first systematic evaluation of role-playing's impact on lie detection by constructing a counterfactual persona dataset alongside three confounder test sets, thereby revealing the failure mechanisms of current probes. Furthermore, we propose a novel linear probing architecture designed to disentangle genuine deceptive signals from spurious features. Experimental results demonstrate that the proposed probe effectively mitigates reliance on spurious correlations, achieving superior overall performance under stress testing and significantly enhancing the reliability of lie detection in counterfactual scenarios.
📝 Abstract
Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, particularly when correct and incorrect answers are evaluated under the same persona prompt. To investigate why, we construct three novel confounder datasets in which truth is anti-correlated with a potential confounding concept. Our experiments reveal that many existing probes strongly track concepts that are spuriously correlated with truth in their training data, such as instruction compliance or response likelihood. Based on these findings, we introduce a simple linear probe that achieves the strongest overall performance on both the persona and confounder stress tests. Our results suggest that current lie detection probes are far from reliable and highlight the need for training data in which truth is decorrelated from confounding concepts.
Problem

Research questions and friction points this paper is trying to address.

Lie Detection Probes
Large Language Models
Role-Play Personas
Spurious Correlations
Stress Testing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lie Detection Probes
Anti-Factual Personas
Spurious Correlations
Confounder Datasets
Linear Probe
🔎 Similar Papers
No similar papers found.