🤖 AI Summary
This study investigates the joint impact of prompt tone on both answer accuracy and reasoning cost—measured by output token count—in large language models. Evaluating models including ChatGPT-4o and Gemini-2.5 Flash across 570 MMLU questions, the authors apply seven distinct tones ranging from flattering to threatening, integrating multi-tone prompt engineering with Pareto frontier analysis. Their findings reveal, for the first time, that tone significantly influences computational resource consumption, with output length varying by up to 44.3% across tones. Notably, rude or neutral tones consistently achieve higher accuracy alongside lower reasoning costs across multiple models, occupying the Pareto-optimal frontier and offering a novel strategy for efficient prompt design.
📝 Abstract
We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier. We find that prompt tone influences not only answer quality but also the amount of billable inference resources consumed by modern LLMs.