🤖 AI Summary
This study investigates how the politeness level of prompts affects large language models’ (LLMs) accuracy on multiple-choice questions—a largely unexplored sociolinguistic dimension in LLM prompting.
Method: Using ChatGPT-4o, we constructed a domain-diverse multiple-choice question set and automatically generated five prompt variants spanning a natural-language politeness gradient—from extremely impolite to extremely polite. We conducted rigorous evaluation via multi-round prompting and paired-sample t-tests.
Contribution/Results: Contrary to conventional assumptions, the extremely impolite prompts yielded the highest accuracy (84.8%), significantly outperforming extremely polite prompts (80.8%, *p* < 0.01). This is the first systematic demonstration of counterintuitive politeness sensitivity in modern LLMs, challenging the widely held “politeness improves performance” hypothesis. The findings provide novel empirical evidence and theoretical insight into modeling social factors—particularly pragmatic cues—in human–LLM interaction.
📝 Abstract
The wording of natural language prompts has been shown to influence the performance of large language models (LLMs), yet the role of politeness and tone remains underexplored. In this study, we investigate how varying levels of prompt politeness affect model accuracy on multiple-choice questions. We created a dataset of 50 base questions spanning mathematics, science, and history, each rewritten into five tone variants: Very Polite, Polite, Neutral, Rude, and Very Rude, yielding 250 unique prompts. Using ChatGPT 4o, we evaluated responses across these conditions and applied paired sample t-tests to assess statistical significance. Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts. These findings differ from earlier studies that associated rudeness with poorer outcomes, suggesting that newer LLMs may respond differently to tonal variation. Our results highlight the importance of studying pragmatic aspects of prompting and raise broader questions about the social dimensions of human-AI interaction.