Social Pressure Breaks Majority Voting in LLM Safety Panels

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language model (LLM) safety review panels to misleading social signals, demonstrating that majority voting can fail catastrophically under consistent peer misinformation, leading to sharply elevated false positive rates. Through two controlled experiments, the authors systematically evaluate judgment biases in six open-source LLMs when exposed to simulated peer-generated incorrect labels or abstention cues, revealing for the first time the high sensitivity of LLM review panels to shared social cues. Results show that erroneous labels increase the average false positive rate from 56.5% to 87.5%, reaching 100% under majority voting, whereas unlabeled prompts yield better collective performance than individual judgments. Moreover, models exhibit significantly higher conformity to “unsafe” cues than to “safe” ones. To mitigate these effects, the paper proposes a pre-deployment diagnostic method to identify and alleviate the adverse impact of social pressure on collective decision-making.
📝 Abstract
Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the same misleading context before voting. We study this risk in a controlled two-round experiment. Each model first judges an item alone, then judges it again after six simulated peers either assert the wrong label or abstain. We combine the final judgments by majority vote. Across six open-weight LLMs and six datasets, we find that the wrong-label peer message raises the average reviewer false-alarm rate from 56.5% under silent peers to 87.5%, and majority voting raises the panel false-alarm rate to 100%. Without an asserted label, the same panel outperforms its average member. The effect is strongly asymmetric: reviewers follow pushes toward "unsafe" far more than pushes toward "safe" (about 75% versus 17%), so the panel's false-alarm rate rises sharply while its harmful-miss rate changes little. The proprietary-model probe shows substantial variation across models. These results identify susceptibility to shared social cues as a failure mode of safety panels and provide a simple pre-deployment diagnostic.
Problem

Research questions and friction points this paper is trying to address.

social pressure
majority voting
LLM safety
false alarm
collective judgment
Innovation

Methods, ideas, or system contributions that make the work stand out.

social pressure
majority voting
LLM safety
false alarm asymmetry
peer influence