Mind Your Tone: Does Tone Alter LLM Performance?

📅 2026-05-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the impact of prompt phrasing variations on the accuracy of large language models (LLMs) on objective multiple-choice questions. By systematically evaluating four prominent models across two datasets containing diverse phrasing variants on subsets of the MMLU benchmark, and complementing the analysis with statistical significance testing, the work reveals that phrasing effects are systematic and highly dependent on both the specific model and subject domain. The findings demonstrate that certain models exhibit substantial accuracy fluctuations in response to phrasing changes, with sensitivity varying across disciplines. Building on these insights, the paper proposes a routing framework that elucidates how phrasing modulates internal reasoning pathways within LLMs, cautioning against the assumption that models are inherently robust to prompt phrasing in practical applications.
📝 Abstract
The use of Large Language Models (LLMs) is proliferating, yet their performance is observed to vary based on prompting styles and tones. In this study, we investigate both whether and how tonal variations in prompts lead to disparate LLM accuracy for objective multiple-choice questions. We use two datasets: a 50-base question dataset with five tone variants and a 570-base question MMLU subset spanning 57 subjects with seven tone variants. Experiments were conducted to evaluate the performance of four cost-efficient, popular LLMs: ChatGPT-4o, ChatGPT-5-nano, Gemini 2.5 Flash, and Gemini 2.5 Flash Lite. Across models, tonal effects are systematic but highly model-dependent. Some models show small, yet statistically significant, shifts, while others exhibit large accuracy swings across tones. Further, we identify subject-level differences in tone sensitivity and present a routing framework to explain how tones may attune internal reasoning modes. Our findings caution users against assuming tone-robust reliability in LLM deployments.
Problem

Research questions and friction points this paper is trying to address.

tone
Large Language Models
prompting
accuracy
LLM performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

tone sensitivity
prompt engineering
large language models
reasoning modes
model robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.